Skip to content
All roles

These are sample roles, not open positions.

The board is not connected to real listings yet. Everything below is placeholder content used to build the page — the companies, the salaries and the reviewers are invented. Nothing here can be applied to.

Cadence-9 · London · Hybrid · Posted

Training Infrastructure Engineer

JAXRayKubernetesNCCL
Compensation
£145k–£190k
Scale you'd work at
>1k accelerators
Discipline
Training infrastructure
Arrangement
Hybrid · Full-time

Our training runs are large enough that hardware failure is a scheduled event rather than an incident, and long enough that a silent stall costs more than a crash. The platform is JAX on Ray over Kubernetes, and the job is keeping utilisation high while making every interruption explain itself.

Much of the work is unglamorous and high-leverage: checkpoint and restore paths that actually resume, collective-communication topology that survives a degraded node, and instrumentation that distinguishes a slow step from a stuck one before a human has to guess.

You will work directly with the research team, whose priorities will change, and whose runs you will occasionally have to stop.

What we’re looking for

  • Operated multi-node distributed training at meaningful scale
  • Debugged an NCCL or collective-communication failure to root cause
  • Comfortable in Kubernetes when it is the problem, not the solution
  • Can reason about checkpoint correctness, not just checkpoint frequency

What this role asks you

Tell us about a distributed training run that failed in a way that took you a long time to understand. What was actually wrong, and what did you change afterwards?

That is the application. No cover letter, no screener, no timed test — Marcus reads your answer.

You would report to Tom Alderidge, Head of Training Platform.