Cadence-9 · London · Hybrid · Posted
Training Infrastructure Engineer
- £145k–£190k
- >1k accelerators
- Training infrastructure
- Hybrid · Full-time
Our training runs are large enough that hardware failure is a scheduled event rather than an incident, and long enough that a silent stall costs more than a crash. The platform is JAX on Ray over Kubernetes, and the job is keeping utilisation high while making every interruption explain itself.
Much of the work is unglamorous and high-leverage: checkpoint and restore paths that actually resume, collective-communication topology that survives a degraded node, and instrumentation that distinguishes a slow step from a stuck one before a human has to guess.
You will work directly with the research team, whose priorities will change, and whose runs you will occasionally have to stop.
What we’re looking for
- Operated multi-node distributed training at meaningful scale
- Debugged an NCCL or collective-communication failure to root cause
- Comfortable in Kubernetes when it is the problem, not the solution
- Can reason about checkpoint correctness, not just checkpoint frequency
Tell us about a distributed training run that failed in a way that took you a long time to understand. What was actually wrong, and what did you change afterwards?
That is the application. No cover letter, no screener, no timed test — Marcus reads your answer.
You would report to Tom Alderidge, Head of Training Platform.