Back to results

Research Staff, Voice AI Foundations

$150,000 – $250,000Remote

  • Engineering
  • United States
  • Full time
  • 2d ago
Newly posted

About the role

As a member of the research staff, you will pioneer the development of Latent Space Models to address fundamental data, scale, and cost challenges in voice AI. You will work on next-generation neural audio codecs, steerable generative models, and multimodal speech-to-speech systems. This role requires an AI-first mindset and the ability to adapt quickly to the rapid pace of AI innovation.

Responsibilities

  • Build next-generation neural audio codecs for extreme, low bit-rate compression and high fidelity reconstruction.
  • Pioneer steerable generative models capable of synthesizing diverse human speech and emotional expression.
  • Develop embedding systems that factorize codec latent space into interpretable dimensions for precise control.
  • Leverage latent recombination to generate synthetic audio data at scale.
  • Train multimodal speech-to-speech systems that produce empathic, human-like responses.
  • Design model architectures, training schemes, and inference algorithms optimized for bare-metal hardware.

Required skills

  • Statistical learning theory
  • Foundation model architectures
  • Multimodal learning
  • Data pipelines
  • Experimental design
  • Speech AI
  • Language AI

Qualifications

  • Strong mathematical foundation in statistical learning theory
  • Deep expertise in foundation model architectures
  • Proven ability to bridge theory and practice
  • Demonstrated ability to build data pipelines for massive datasets

About the Company

Deepgram is a platform for the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech, and building production-grade voice agents. The company's voice-native foundation models are used by over 200,000 developers and 1,300 organizations.

Research Staff, Voice AI Foundations at Deepgram · Grasshire