Verge Robotics · Zürich · Posted
RL Engineer, Manipulation
- €155k–€205k
- Multi-node
- RL & agents
- On-site · Full-time
We train manipulation policies in simulation and deploy them onto real arms, which means most of the interesting work happens in the gap between the two. Reward specification, domain randomisation, and the long tail of ways a policy can technically succeed while doing something absurd.
The role is hands-on with both halves: you will write the reward, watch the robot exploit it, and rewrite it. Onsite in Zürich because the hardware is here and the feedback loop matters more than the flexibility.
We are looking for someone who has watched a policy do something clever and wrong, and enjoyed it.
What we’re looking for
- Trained RL policies that ran on physical hardware
- Have hit — and diagnosed — a reward-hacking failure
- Comfortable with MuJoCo or an equivalent contact-rich simulator
- Can work alongside hardware and controls engineers
Tell us about a policy that learned to exploit your reward function. What did it discover, and how did you fix the specification?
That is the application. No cover letter, no screener, no timed test — Sofia reads your answer.
You would report to Lukas Brenner, Head of Manipulation.