Skip to content
All roles

These are sample roles, not open positions.

The board is not connected to real listings yet. Everything below is placeholder content used to build the page — the companies, the salaries and the reviewers are invented. Nothing here can be applied to.

Northrail · Remote (US) · Posted

Evals Lead, Post-training

PythonPyTorchWeights & Biases
Compensation
$200k–$265k
Scale you'd work at
Single-node
Discipline
Evals & measurement
Arrangement
Remote · Full-time

Every post-training decision here routes through an eval suite, which means the suite is load-bearing and its failures are expensive. We have caught contamination late more than once, and the cost of that is measured in weeks.

You would own the measurement layer: what we test, how we know a benchmark has leaked, and how we decide a model is better rather than merely different. That includes the unfashionable half — provenance, dataset hygiene, and being the person who says a result is not real yet.

The hard part is not writing evals. It is building the ones that stay honest as the team learns to optimise against them.

What we’re looking for

  • Built an eval suite that a team actually shipped decisions on
  • Have found a contamination or leakage problem, and can describe how
  • Comfortable arguing a result is not significant, to people who want it to be
  • Strong statistical intuition for small-sample comparisons

What this role asks you

Describe an evaluation you built or inherited that turned out to be measuring the wrong thing. How did you notice, and what did you do about it?

That is the application. No cover letter, no screener, no timed test — Wei reads your answer.

You would report to Dana Whitfield, VP Research.