You make performance claims useful to the person buying the computer, by measuring what the system accomplishes, what it costs, and how reliably it does it. This role exists partly to keep us honest: you are the person who will identify our own cherry-picked workload and say so.
Open for applications. Starts at: Pilot expansion.
We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.
Where
In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.
The work
Build representative task suites covering inference, retrieval, training and integrations. Record model version, quality, context, precision, memory and device conditions. Measure latency distributions, wall energy and completed-task cost alongside raw throughput.
The milestone
In your first 90 days, publish a repeatable baseline and a comparison that another engineer can reproduce on the same configuration.
Required
Nice to have
Evidence
Bring experimental design, systems measurement and clear technical writing. Be able to identify cherry-picked workloads and distinguish measurement noise from a meaningful change.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Design a fair comparison between two devices when their fastest available model configurations have different quality.
The package
Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.
Apply
One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.