You make research results and model releases traceable to the data and decisions that produced them. Most irreproducibility is a data-lineage problem rather than a modelling one, and this role is the infrastructure that makes learning responsibly possible rather than aspirational.
Open for applications. Starts at: Pilot expansion.
We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.
Where
In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.
The work
Version datasets, model artifacts, experiment configurations and evaluation results. Build provenance, quality checks and reproducible pipelines. Separate household data from shared research datasets, and coordinate deletion and retention behavior with privacy engineers.
The milestone
In your first 90 days, deliver a reproducible experiment pipeline with dataset lineage and release checks that block unapproved inputs.
Required
Nice to have
Evidence
Bring experience with data systems, ML operations or scientific computing. Show how you have detected contamination, data drift or irreproducible results.
Evidence, not credentials. We are describing work you can point at, in whatever form it exists.
The exercise
Trace a model-quality regression after a dataset update and explain which artifacts are necessary to reproduce it.
The package
Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.
Apply
One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.