Speech & Audio jobs

Voice became a serious product surface again once models could listen and talk without embarrassment, and hiring followed: ASR, text-to-speech, speaker diarization, and the real-time plumbing that keeps a conversation feeling instant. Whisper made transcription a commodity; making it conversational did not get any easier.

Stacks run from Whisper fine-tunes and Kaldi holdouts to modern streaming architectures, with WebRTC knowledge suddenly valuable again. Companies building voice agents care about barge-in handling and latency budgets as much as word error rate. It's a smaller field than computer vision, which cuts both ways: fewer postings, far fewer qualified applicants.

26open roles right now

see all with filters