You turn models into a dependable local service, running responsively on a computer the person owns. Local inference is where the sovereignty claim is either true or marketing, and the constraint is unforgiving: fixed memory, no autoscaling, and a user who notices every pause.
Open for applications. Starts at: First team.
We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.
Where
In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.
The work
Build model serving, memory management, quantization evaluation, batching and scheduling. Support the chosen accelerator backends and measure startup, time to first token, sustained throughput and quality. Keep interactive tasks responsive while background work uses spare capacity.
The milestone
In your first 90 days, ship a local inference service with a reproducible performance and quality report for the reference device.
Required
Nice to have
Evidence
Bring strong performance engineering and practical experience serving machine-learning models. Understand memory capacity, bandwidth, context length and why theoretical compute figures do not predict user experience.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Diagnose an inference slowdown as context grows and propose an improvement without quietly reducing output quality.
The package
Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.
Apply
One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.