Head of Inference
Series C AI lab, 120 people
Own model serving for a lab moving from research demos to paid API traffic at 40,000 requests a second.
Practice
Inference & infrastructure
Location
London
Compensation
£260k–£340k + equity
Progress
Week 5 of 9
01 — Mapped
4,812
02 — Called by a partner
212
03 — Met for two hours
38
04 — Shortlisted
6
■ The brief
Partner on this search
Priya Raman
Every conversation is confidential. We never share a name without permission.
Start a conversation
→
The brief
Their serving stack was built by researchers for demos. It now carries paying customers, and the CTO is its on-call engineer. The hire owns latency, cost per token and the team of six that keeps both honest.
Who we’re looking for
Someone who has run inference for a product people pay for: batching, KV-cache strategy, speculative decoding, the unglamorous work of GPU utilisation. Frontier-lab or hyperscaler serving teams, or a startup that grew up fast.
How the search runs
A named partner runs the search from brief to offer. We map the market in the first two weeks, make every first approach ourselves, and share the funnel with you every Friday: who we called, who said no, and why.
Interested, or know someone who should be? Write to the partner on this search. Every conversation is confidential, and we never send a CV without your permission.
More
■ Other searches
Open and closed searches across our four practices.
#0412
Head of Inference
Week 5 of 9
#0409
Quant Researcher, Futures
Week 3 of 8
#0407
Founding ML Engineer
Week 6 of 8
■ Know someone?
Refer the person you’d hire yourself.
If they’re placed through a referral, we give the fee equivalent of a month’s work to the charity of your choice. Your name stays out of it unless you want it in.
