Alignment Research jobs
Alignment research is the work of making capable models want what we meant: reward modeling, RLHF and constitutional-style methods, scalable oversight, and the theory of what training actually optimizes. Most seats sit at frontier labs, with a smaller ring at nonprofits and university-adjacent groups.
The field prizes people who can turn a philosophical worry into an experiment with a metric by Friday. Empirical work dominates hiring — training reward models, building oversight protocols, studying deceptive behavior in controlled settings. Openings are few and applications many, but published independent work still genuinely moves candidates.
8open roles right now
- Researcher, Recursive Self-Improvement SafetyOpenAIOn-site · Senior · AI Safety$295k–$445k
- Researcher, Alignment CoT MonitorabilityOpenAIHybrid · Senior · AI Safety$250k–$445k
- Anthropic Fellows Program, AI SafetyAnthropicRemote (US, UK, Canada) · Intern · AI Safety$62k
- Program Manager, AlignmentOpenAIHybrid · Mid-level · AI Safety$162k–$240k
- Research Engineer, AI Safety & AlignmentCharacter.AIOn-site · Senior · AI Safety$225k–$400k
- Research Engineer / Scientist, AlignmentAnthropicHybrid · Mid-level · AI Safety$350k–$500k
- [Expression of Interest] Research Engineer / Scientist, Alignment - LondonAnthropicRemote (UK) · Mid-level · AI Safety£260k–£370k
- Researcher, AlignmentOpenAIRemote (US, UK) · Senior · AI Safety$250k–$445k