Training Infrastructure jobs

Training infrastructure teams build the machinery under the runs: distributed frameworks, checkpointing that survives node failures, data loaders that keep thousands of GPUs from starving, and the observability to explain a slow step time. At sufficient scale hardware fails daily, which makes reliability engineering its own discipline here.

The vocabulary is FSDP, Megatron, DeepSpeed, and JAX's sharding machinery, plus storage tuned for sequential reads at absurd throughput. Debugging spans the whole stack, from NCCL timeouts to a flaky top-of-rack switch. Labs hire this continuously, and strong systems engineers without ML backgrounds cross over here regularly.

137open roles right now

see all with filters