Alignment Research jobs
Alignment research is the work of making capable models want what we meant: reward modeling, RLHF and constitutional-style methods, scalable oversight, and the theory of what training actually optimizes. Most seats sit at frontier labs, with a smaller ring at nonprofits and university-adjacent groups.
The field prizes people who can turn a philosophical worry into an experiment with a metric by Friday. Empirical work dominates hiring — training reward models, building oversight protocols, studying deceptive behavior in controlled settings. Openings are few and applications many, but published independent work still genuinely moves candidates.
12open roles right now
- Model Policy Manager, Agentic SafetyOpenAIHybrid · Senior · AI Safety$207k–$335k
- Research Program Manager, GovernanceOpenAIOn-site · Senior · AI Safety$239k–$328k
- Researcher, Agent Safety, Oversight and System MitigationsOpenAIHybrid · Senior · AI Safety$380k–$500k
- Researcher, Alignment InterpretabilityOpenAIOn-site · Senior · AI Safety$295k–$500k
- Researcher, Frontier Risk MitigationsOpenAIOn-site · Senior · AI Safety$380k–$500k
- Researcher, Recursive Self-Improvement SafetyOpenAIOn-site · Senior · AI Safety$380k–$500k
- Researcher, Alignment CoT MonitorabilityOpenAIHybrid · Senior · AI Safety$295k–$500k
- Anthropic Fellows Program, AI Safety & SecurityAnthropicRemote (US, UK, Canada) · Intern · AI Safety$62k
- Research Engineer, AI Safety & AlignmentCharacter.AIOn-site · Senior · AI Safety$225k–$400k
- Research Engineer / Scientist, AlignmentAnthropicHybrid · Mid-level · AI Safety$350k–$500k
- [Expression of Interest] Research Engineer / Scientist, Alignment - LondonAnthropicRemote (UK) · Mid-level · AI Safety£260k–£370k
- Researcher, AlignmentOpenAIRemote (US, UK) · Senior · AI Safety$295k–$500k