Skip to content
All roles

These are sample roles, not open positions.

The board is not connected to real listings yet. Everything below is placeholder content used to build the page — the companies, the salaries and the reviewers are invented. Nothing here can be applied to.

Trellis Health · Remote (US) · Boston · Posted

ML Data Engineer

SparkdbtAirflowDuckDB
Compensation
$160k–$205k
Scale you'd work at
Petabyte-scale
Discipline
Data engineering
Arrangement
Remote · Full-time

Clinical data is messy in ways that matter: it arrives late, it arrives twice, and it arrives with the label you wanted derived from something that would not have been available at prediction time. Our pipelines are the layer that catches that before a model does.

You would own ingestion through feature production — Spark and dbt over a petabyte-scale store, orchestrated with Airflow — with a strong emphasis on lineage and on making leakage structurally difficult rather than merely discouraged.

This is a role where the most valuable thing you build may be the test that stops a pipeline.

What we’re looking for

  • Built production data pipelines feeding ML training
  • Have caught a target-leakage bug, and can explain the mechanism
  • Fluent in Spark and SQL at a scale where the query plan matters
  • Care about lineage and reproducibility as first-class concerns

What this role asks you

Describe a data pipeline you built where correctness was hard to guarantee. What could have gone silently wrong, and what did you put in place to catch it?

That is the application. No cover letter, no screener, no timed test — Daniel reads your answer.

You would report to Renee Castellanos, Director of Data Platform.