Trellis Health · Remote (US) · Boston · Posted
ML Data Engineer
- $160k–$205k
- Petabyte-scale
- Data engineering
- Remote · Full-time
Clinical data is messy in ways that matter: it arrives late, it arrives twice, and it arrives with the label you wanted derived from something that would not have been available at prediction time. Our pipelines are the layer that catches that before a model does.
You would own ingestion through feature production — Spark and dbt over a petabyte-scale store, orchestrated with Airflow — with a strong emphasis on lineage and on making leakage structurally difficult rather than merely discouraged.
This is a role where the most valuable thing you build may be the test that stops a pipeline.
What we’re looking for
- Built production data pipelines feeding ML training
- Have caught a target-leakage bug, and can explain the mechanism
- Fluent in Spark and SQL at a scale where the query plan matters
- Care about lineage and reproducibility as first-class concerns
Describe a data pipeline you built where correctness was hard to guarantee. What could have gone silently wrong, and what did you put in place to catch it?
That is the application. No cover letter, no screener, no timed test — Daniel reads your answer.
You would report to Renee Castellanos, Director of Data Platform.