Back to results

Elevenlabs

Visit website

Research Engineer - Inference

Salary not disclosedRemote

  • Engineering
  • Bulgaria
  • Full time
  • Today
Newly posted

About the role

ElevenLabs is seeking a Research Engineer to join their research team, focusing on the deployment and optimization of frontier AI models in production. The role involves owning the systems that transition research breakthroughs into real-time products used by millions of users. You will be responsible for ensuring models are served fast, reliably, and at scale.

Responsibilities

  • Deploy state-of-the-art models to production and own the path from research checkpoint to serving infrastructure.
  • Optimize inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
  • Build and tune high-performance serving systems for real-time, streaming workloads.
  • Create tooling and infrastructure that enables researchers to ship new models to production quickly, safely, and with confidence in their performance characteristics.

Required skills

  • ML model deployment
  • GPU programming
  • Inference optimization
  • CUDA
  • Triton
  • TensorRT
  • vLLM
  • SGLang

Qualifications

  • Equivalent practical experience

Benefits

  • Annual professional development stipend
  • Annual social travel stipend
  • Annual company offsite
  • Monthly co-working stipend

About the Company

ElevenLabs is an AI research and product company that launched in January 2023 with the first human-like AI voice model. They serve millions of users and thousands of businesses, with platforms spanning voice agents, creative tools, and developer APIs.