The GPUs don't run themselves.

AI Infrastructure jobs

AI infrastructure is where the GPUs actually live: cluster scheduling, high-speed networking, storage fast enough to feed a training run, and inference serving that stays up on launch day. Titles include ML Platform Engineer, Inference Engineer, HPC Engineer, and the occasional SRE who wandered over and stayed.

The stack is Kubernetes with Ray or Slurm for orchestration, Terraform underneath, and CUDA, NCCL, and TensorRT-LLM when things get low-level — knowing why a quantized model misbehaves is a genuine superpower here. Demand comfortably exceeds supply, and because the clusters are remote anyway, many of the jobs are too. Pay is strong and posted more often than you'd guess.

— Specialisms

  1. 01GPU Infrastructure173
  2. 02Inference & Serving176
  3. 03Training Infrastructure117
  4. 04ML Platform251
  5. 05Data Infrastructure75
  6. 06AI Security19

818open roles right now

see all with filters