You make each unit of compute do more useful work, optimising the kernels and compilation paths that matter to real workloads rather than to a paper. Mastery of every backend is not expected; depth in one and the judgement to find the bottleneck is.
Open for applications. Starts at: Pilot expansion.
We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.
Where
In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.
The work
Profile operators, memory traffic and scheduling; implement and validate kernels on selected backends. Work with numerical and model researchers on precision choices. Maintain portable abstractions where useful while documenting backend-specific behavior and limits.
The milestone
In your first 90 days, deliver one measured end-to-end improvement with correctness tests, hardware details and a reproducible comparison.
Required
Nice to have
Evidence
Bring GPU programming, compiler or numerical computing expertise. Relevant experience may include CUDA, Metal, LLVM, MLIR, Triton or other accelerator toolchains; mastery of every backend is unnecessary.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Show how a faster isolated kernel could still make the complete application slower, and design the correct benchmark.
The package
Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.
Apply
One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.