The GPUs don't run themselves.

AI Infrastructure jobs

AI infrastructure is where the GPUs actually live: cluster scheduling, high-speed networking, storage fast enough to feed a training run, and inference serving that stays up on launch day. Titles include ML Platform Engineer, Inference Engineer, HPC Engineer, and the occasional SRE who wandered over and stayed.

The stack is Kubernetes with Ray or Slurm for orchestration, Terraform underneath, and CUDA, NCCL, and TensorRT-LLM when things get low-level — knowing why a quantized model misbehaves is a genuine superpower here. Demand comfortably exceeds supply, and because the clusters are remote anyway, many of the jobs are too. Pay is strong and posted more often than you'd guess.

— Specialisms

  1. 01GPU Infrastructure100
  2. 02Inference & Serving156
  3. 03Training Infrastructure137
  4. 04ML Platform184
  5. 05Data Infrastructure60
  6. 06AI Security16

659open roles right now

see all with filters