hussh
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
Get Agent One
Agent One
Products
PuppyTagShop
For Business
For Advisors & RIAFor BrandsFor Agent BuildersPartner with Hussh
Blog
LatestProduct UpdatesResearchFounder’s Notes
Glossary
A-ZConsent & PrivacyAgents & AIData Ownership
Company
Our StoryTeamCareersPressContact
AI infrastructure · H27

AI Inference Systems Engineer

You turn models into a dependable local service, running responsively on a computer the person owns. Local inference is where the sovereignty claim is either true or marketing, and the constraint is unforgiving: fixed memory, no autoscaling, and a user who notices every pause.

Apply for this roleAll roles

Open for applications. Starts at: First team.

We are taking applications for this role now and building the pipeline for it. The stage above is when the work itself is expected to begin, which is something you deserve to know before you apply rather than after. It is context, not a gate.

Where

Kirkland GarageUAE Garage

In the office together five days a week, in one of our garages, and remote-friendly around your family, arranged one person at a time. We hire across the United States 🇺🇸, India 🇮🇳 and the UAE 🇦🇪.

The work

What this person actually does

Build model serving, memory management, quantization evaluation, batching and scheduling. Support the chosen accelerator backends and measure startup, time to first token, sustained throughput and quality. Keep interactive tasks responsive while background work uses spare capacity.

The milestone

What it looks like when it is working

In your first 90 days, ship a local inference service with a reproducible performance and quality report for the reference device.

Required

What we would not hire without

  • ▸Strong performance engineering plus practical experience serving machine-learning models
  • ▸You reason about memory capacity, bandwidth and context length together rather than in isolation
  • ▸You know why theoretical compute figures do not predict real throughput, and measure instead
  • ▸You optimise against a latency budget somebody set for a user, not for a benchmark

Nice to have

What would be a bonus, not a gate

  • ▸Quantisation, speculative decoding, or KV-cache optimisation
  • ▸Apple silicon or consumer GPU targets
  • ▸You have shipped an inference stack somebody else ran

Evidence

What would show us you can do it

Bring strong performance engineering and practical experience serving machine-learning models. Understand memory capacity, bandwidth, context length and why theoretical compute figures do not predict user experience.

Evidence, not credentials. We are describing work you can point at, in whatever form it exists.

The exercise

How we would look at it together

Diagnose an inference slowdown as context grows and propose an improvement without quietly reducing output quality.

The package

What comes with the job

  • ▸Stock options for every full-time teammate, four-year vesting with a one-year cliff
  • ▸Annual performance bonus, or on-target earnings with uncapped commission for customer-facing roles
  • ▸Medical, dental and vision for you and your family, plus life and disability cover, on the highest plan tier available to us
  • ▸A 401(k) with company matching
  • ▸Pay reviewed every year and on promotion, benchmarked to your role and market
  • ▸A budget of AI tokens of your own
  • ▸Gym membership, and retailer discounts redeemed through our benefits app
  • ▸Remote-friendly around your family, arranged one person at a time
  • ▸$1,000 plus $10,000 in equity for a referral we hire who stays a year

Indicative pay ranges by market and level are on the compensation page. Plan numbers are confirmed in your offer letter.

Apply

Apply for AI Inference Systems Engineer

One short form. A person reads every application and you hear back either way. You will get your own link to check where things stand, and you can withdraw or delete your application from it at any time, without an account.

Apply for this job

* indicates a required field

AI Inference Systems Engineer

Resume *

Any one of these. If you have not got a PDF to hand, paste the text - it is not a lesser way to apply.

That is everything we need. The rest is optional, and it helps.

Where your work lives

Any of these, none of these. Paste a link and we will look.

In a hundred words or so, the piece of work you are most proud of.

You get a reference number straight away.
← All 72 roles in the catalog