Halide Labs · Remote (EU / US) · Posted
Inference Engineer, vLLM & CUDA
- $180k–$240k
- 8–64 GPUs
- Inference & serving
- Remote · Full-time
We serve a 70B model family to production traffic with a p99 budget that has not moved in a year, while the traffic behind it has tripled. The serving path is vLLM with a set of custom Triton kernels for the attention hot loop, and the next year of work is making that path cheaper without making it slower.
You would own the batching and cache layer end to end: continuous batching policy, KV-cache eviction, prefix reuse across a request population that is far more repetitive than it looks. There is an existing benchmark harness and it is honest — it will happily tell you your change made things worse.
This is a role for someone who profiles before optimising, and who has opinions about what a p99 number is actually measuring.
What we’re looking for
- Shipped an inference path to production traffic, not just a benchmark
- Comfortable reading and writing CUDA or Triton kernels
- Have measured — and can explain — a real latency regression
- Familiar with paged or block-based KV-cache designs
Describe a time you profiled an inference bottleneck. What did you measure, what did you change, and how did you know the change worked?
That is the application. No cover letter, no screener, no timed test — Ines reads your answer.
You would report to Priya Raman, Director of Inference.