Multimodal Research jobs

Multimodal research is where the modalities stop being separate fields: vision-language models, audio understanding, video reasoning, and any-to-any architectures that treat everything as tokens. Labs want people who can design the training mix across modalities and debug why adding images made the text worse.

The background that lands these roles is usually depth in one modality plus fluency in the language-model stack — a computer vision PhD who knows post-training, say. Data pipelines are half the challenge, because paired multimodal data is scarcer and messier than text. Geography follows the labs, and the pay bands stay unpublished, as with the rest of frontier research.

17open roles right now

see all with filters