Skip to content
All roles

These are sample roles, not open positions.

The board is not connected to real listings yet. Everything below is placeholder content used to build the page — the companies, the salaries and the reviewers are invented. Nothing here can be applied to.

Halide Labs · Remote (EU / US) · Posted

Inference Engineer, vLLM & CUDA

PyTorchTritonCUDAvLLM
Compensation
$180k–$240k
Scale you'd work at
8–64 GPUs
Discipline
Inference & serving
Arrangement
Remote · Full-time

We serve a 70B model family to production traffic with a p99 budget that has not moved in a year, while the traffic behind it has tripled. The serving path is vLLM with a set of custom Triton kernels for the attention hot loop, and the next year of work is making that path cheaper without making it slower.

You would own the batching and cache layer end to end: continuous batching policy, KV-cache eviction, prefix reuse across a request population that is far more repetitive than it looks. There is an existing benchmark harness and it is honest — it will happily tell you your change made things worse.

This is a role for someone who profiles before optimising, and who has opinions about what a p99 number is actually measuring.

What we’re looking for

  • Shipped an inference path to production traffic, not just a benchmark
  • Comfortable reading and writing CUDA or Triton kernels
  • Have measured — and can explain — a real latency regression
  • Familiar with paged or block-based KV-cache designs

What this role asks you

Describe a time you profiled an inference bottleneck. What did you measure, what did you change, and how did you know the change worked?

That is the application. No cover letter, no screener, no timed test — Ines reads your answer.

You would report to Priya Raman, Director of Inference.