Data Infrastructure jobs

Data infrastructure at AI scale is a storage and throughput problem first: petabyte object stores, streaming ingestion, metadata catalogs, and a query layer that keeps analysts and training jobs fed without either starving the other. Training-run reads are sequential, enormous, and impatient, which bends most warehouse-era assumptions.

Kafka or its successors handle ingestion, Iceberg or Delta sit over object storage, and Spark and Ray do the processing, with GPU-adjacent caching layers a growing specialty. Cost engineering is a first-class duty — data platforms quietly become the second-biggest line item after compute. Deep distributed-systems experience counts for more here than any ML pedigree.

60open roles right now

see all with filters