Reinforcement Learning jobs

Reinforcement learning came back from the wilderness the moment reasoning models made RL-on-language the hottest recipe in the field. The jobs split between RL for LLMs — reward modeling, verifiable rewards, environment design — and the classical control end that never left robotics and operations research.

You'll want fluency in policy-gradient methods and their failure modes, plus enough engineering to build environments that don't leak reward. Code and math environments, agentic benchmarks, and reward-hacking forensics dominate current postings. Genuine RL experience is a seller's market, and interviews go deep on the algorithmic details.

23open roles right now

see all with filters