Alignment Research jobs

Alignment research is the work of making capable models want what we meant: reward modeling, RLHF and constitutional-style methods, scalable oversight, and the theory of what training actually optimizes. Most seats sit at frontier labs, with a smaller ring at nonprofits and university-adjacent groups.

The field prizes people who can turn a philosophical worry into an experiment with a metric by Friday. Empirical work dominates hiring — training reward models, building oversight protocols, studying deceptive behavior in controlled settings. Openings are few and applications many, but published independent work still genuinely moves candidates.

8open roles right now

see all with filters