hussh
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
Get Agent One
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
AI research · H38

AI Evaluation and Reliability Scientist

You show when a private agent deserves a user's trust, building the evidence about task completion, failure, and human oversight. Two failure modes govern this role: benchmark contamination, and quietly substituting a model judge for ground truth because it is cheaper.

Apply for this roleAll roles

Open for applications. Starts at: Pilot expansion.

We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.

Where

Kirkland GarageUAE Garage

In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.

The work

What this person actually does

Design evaluations for real workflows, adversarial inputs and distribution shifts. Use held-out tasks, blinded review where appropriate and calibrated human evaluation. Measure false completion, unauthorized actions, appropriate escalation and sustained usefulness over time.

The milestone

What it looks like when it is working

In your first 90 days, deliver a versioned evaluation suite and a release report that makes strengths and unresolved failures easy to inspect.

Required

What we would not hire without

  • ▸Experimental design, statistics, and hands-on AI evaluation
  • ▸You prevent benchmark contamination deliberately, with a method you can describe
  • ▸You do not substitute a model judge for ground truth, and can explain when a judge is admissible
  • ▸You report failure rates as plainly as success rates

Nice to have

What would be a bonus, not a gate

  • ▸Red-teaming or adversarial evaluation
  • ▸Human evaluation study design
  • ▸You have built an eval suite that changed a shipping decision

Evidence

What would show us you can do it

Bring experimental design, statistics and hands-on AI evaluation. Show how you prevent benchmark contamination and avoid substituting a model judge for ground truth.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Design an evaluation that detects an agent becoming more persuasive without becoming more accurate or reliable.

The package

What comes with the job

  • ▸Stock options for every full-time teammate, four-year vesting with a one-year cliff
  • ▸Annual performance bonus, or on-target earnings with uncapped commission for customer-facing roles
  • ▸Medical, dental and vision for you and your family, plus life and disability cover, on the highest plan tier available to us
  • ▸A 401(k) with company matching
  • ▸Pay reviewed every year and on promotion, benchmarked to your role and market
  • ▸A budget of AI tokens of your own
  • ▸Gym membership, and retailer discounts redeemed through our benefits app
  • ▸Remote-friendly around your family, arranged one person at a time
  • ▸$1,000 plus $10,000 in equity for a referral we hire who stays a year

Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.

Apply

Apply for AI Evaluation and Reliability Scientist

One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.

Apply for this job

* indicates a required field

AI Evaluation and Reliability Scientist

Resume *

Any one of these. If you have not got a PDF to hand, paste the text - it is not a lesser way to apply.

That is everything we need. The rest is optional, and it helps.

Where your work lives

Any of these, none of these. Paste a link and we will look.

In a hundred words or so, the piece of work you are most proud of.

You get a reference number straight away.
← All 72 roles in the catalog