Northrail · Remote (US) · Posted
Evals Lead, Post-training
- $200k–$265k
- Single-node
- Evals & measurement
- Remote · Full-time
Every post-training decision here routes through an eval suite, which means the suite is load-bearing and its failures are expensive. We have caught contamination late more than once, and the cost of that is measured in weeks.
You would own the measurement layer: what we test, how we know a benchmark has leaked, and how we decide a model is better rather than merely different. That includes the unfashionable half — provenance, dataset hygiene, and being the person who says a result is not real yet.
The hard part is not writing evals. It is building the ones that stay honest as the team learns to optimise against them.
What we’re looking for
- Built an eval suite that a team actually shipped decisions on
- Have found a contamination or leakage problem, and can describe how
- Comfortable arguing a result is not significant, to people who want it to be
- Strong statistical intuition for small-sample comparisons
Describe an evaluation you built or inherited that turned out to be measuring the wrong thing. How did you notice, and what did you do about it?
That is the application. No cover letter, no screener, no timed test — Wei reads your answer.
You would report to Dana Whitfield, VP Research.