Interpretability jobs

Interpretability grew from blog posts into funded teams trying to explain what happens inside a transformer — circuits, features, sparse autoencoders, and probes that ride along during training. Some teams sit inside safety orgs, others serve product needs like debugging refusals or steering behavior, and the job differs accordingly.

The days are empirical and code-heavy: training SAEs, running activation patching, and building tooling to look at millions of features without going blind. A strong candidate reads like a scientist and commits like an engineer. The field is small enough that hiring managers all know each other, which makes public artifacts — posts, tools, replications — unusually effective applications.

6open roles right now

see all with filters