Inference & Serving jobs

Inference engineering is where model quality meets a latency budget: serving stacks that stream tokens fast enough to feel alive, batch aggressively enough to be affordable, and stay up while a launch multiplies traffic twenty-fold. The economics are unforgiving — milliseconds and GPU-hours both convert directly into money.

vLLM, SGLang, and TensorRT-LLM dominate the serving layer, with speculative decoding, KV-cache management, and quantization as the standard levers. Kernel fluency in CUDA or Triton separates senior candidates from the queue. This may be the single most in-demand specialty on the board right now, and offers reflect it.

156open roles right now

see all with filters