You show when a private agent deserves a user's trust, building the evidence about task completion, failure, and human oversight. Two failure modes govern this role: benchmark contamination, and quietly substituting a model judge for ground truth because it is cheaper.
Open for applications. Starts at: Pilot expansion.
We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.
Where
In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.
The work
Design evaluations for real workflows, adversarial inputs and distribution shifts. Use held-out tasks, blinded review where appropriate and calibrated human evaluation. Measure false completion, unauthorized actions, appropriate escalation and sustained usefulness over time.
The milestone
In your first 90 days, deliver a versioned evaluation suite and a release report that makes strengths and unresolved failures easy to inspect.
Required
Nice to have
Evidence
Bring experimental design, statistics and hands-on AI evaluation. Show how you prevent benchmark contamination and avoid substituting a model judge for ground truth.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Design an evaluation that detects an agent becoming more persuasive without becoming more accurate or reliable.
The package
Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.
Apply
One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.