hussh
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
Get Agent One
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
AI research · H36

Speech and Multimodal Intelligence Scientist

You make personal computing reachable through the ways people naturally communicate: speech, vision, and whatever else somebody brings to it, under the constraints of a device rather than a data centre. The recurring trap in this area is fluency, which is now easy to produce and easy to mistake for understanding. Telling those apart, and evaluating on the input people really give rather than the input that demos well, is the work.

Apply for this roleAll roles

Open for applications. Starts at: Research program.

We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.

Where

Kirkland GarageUAE Garage

In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.

The work

What this person actually does

Develop or adapt models for speech, documents, images and grounded interaction. Evaluate diverse accents, environments and accessibility needs with appropriate participant consent. Measure local performance and make capture, retention and activation behavior understandable.

The milestone

What it looks like when it is working

In your first 90 days, deliver a bounded multimodal capability with evaluation across realistic conditions and a clear record of its limitations.

Required

What we would not hire without

  • ▸Research or advanced engineering in speech, vision, or multimodal learning
  • ▸You can show how you distinguish apparent fluency from correct understanding
  • ▸You work within real device constraints rather than assuming a server
  • ▸You evaluate on realistic input, including accents, noise, and bad lighting

Nice to have

What would be a bonus, not a gate

  • ▸On-device speech recognition or synthesis
  • ▸Accessibility-driven design
  • ▸Low-resource languages

Evidence

What would show us you can do it

Bring research or advanced engineering in speech, vision or multimodal learning. Show how you distinguish apparent fluency from correct understanding.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Design an evaluation for a voice-controlled task in a shared room where only one person has authorized the action.

The package

What comes with the job

  • ▸Stock options for every full-time teammate, four-year vesting with a one-year cliff
  • ▸Annual performance bonus, or on-target earnings with uncapped commission for customer-facing roles
  • ▸Medical, dental and vision for you and your family, plus life and disability cover, on the highest plan tier available to us
  • ▸A 401(k) with company matching
  • ▸Pay reviewed every year and on promotion, benchmarked to your role and market
  • ▸A budget of AI tokens of your own
  • ▸Gym membership, and retailer discounts redeemed through our benefits app
  • ▸Remote-friendly around your family, arranged one person at a time
  • ▸$1,000 plus $10,000 in equity for a referral we hire who stays a year

Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.

Apply

Apply for Speech and Multimodal Intelligence Scientist

One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.

Apply for this job

* indicates a required field

Speech and Multimodal Intelligence Scientist

Resume *

Any one of these. If you have not got a PDF to hand, paste the text - it is not a lesser way to apply.

That is everything we need. The rest is optional, and it helps.

Where your work lives

Any of these, none of these. Paste a link and we will look.

In a hundred words or so, the piece of work you are most proud of.

You get a reference number straight away.
← All 72 roles in the catalog