Head of Inference

Series C AI lab, 120 people

Own model serving for a lab moving from research demos to paid API traffic at 40,000 requests a second.

Practice

Inference & infrastructure

Location

London

Compensation

£260k–£340k + equity

Progress

Week 5 of 9

01 — Mapped

4,812

02 — Called by a partner

212

03 — Met for two hours

38

04 — Shortlisted

6

■ The brief

Partner on this search

Priya Raman

Every conversation is confidential. We never share a name without permission.

Start a conversation

→

The brief

Their serving stack was built by researchers for demos. It now carries paying customers, and the CTO is its on-call engineer. The hire owns latency, cost per token and the team of six that keeps both honest.

Who we’re looking for

Someone who has run inference for a product people pay for: batching, KV-cache strategy, speculative decoding, the unglamorous work of GPU utilisation. Frontier-lab or hyperscaler serving teams, or a startup that grew up fast.

How the search runs

A named partner runs the search from brief to offer. We map the market in the first two weeks, make every first approach ourselves, and share the funnel with you every Friday: who we called, who said no, and why.

Interested, or know someone who should be? Write to the partner on this search. Every conversation is confidential, and we never send a CV without your permission.

More

■ Other searches

Open and closed searches across our four practices.

■ Know someone?

Refer the person you’d hire yourself.

If they’re placed through a referral, we give the fee equivalent of a month’s work to the charity of your choice. Your name stays out of it unless you want it in.

Create a free website with Framer, the website builder loved by startups, designers and agencies.