Deepgram
Visit websiteResearch Staff, Voice AI Foundations
$150,000 – $250,000Remote
- Engineering
- United States
- Full time
- 2d ago
Newly posted
About the role
As a member of the research staff, you will pioneer the development of Latent Space Models to address fundamental data, scale, and cost challenges in voice AI. You will work on next-generation neural audio codecs, steerable generative models, and multimodal speech-to-speech systems. This role requires an AI-first mindset and the ability to adapt quickly to the rapid pace of AI innovation.
Responsibilities
- Build next-generation neural audio codecs for extreme, low bit-rate compression and high fidelity reconstruction.
- Pioneer steerable generative models capable of synthesizing diverse human speech and emotional expression.
- Develop embedding systems that factorize codec latent space into interpretable dimensions for precise control.
- Leverage latent recombination to generate synthetic audio data at scale.
- Train multimodal speech-to-speech systems that produce empathic, human-like responses.
- Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware.
Required skills
- Statistical learning theory
- Foundation model architectures
- Multimodal learning
- Data pipelines
- Experimental design
- Speech AI
- Language AI
Qualifications
- Strong mathematical foundation in statistical learning theory
- Deep expertise in foundation model architectures
- Proven ability to bridge theory and practice
- Demonstrated ability to build data pipelines for massive datasets
About the Company
Deepgram is a platform for the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech, and building production-grade voice agents. The company's voice-native foundation models are used by over 200,000 developers and 1,300 organizations.