Elevenlabs
Visit websiteResearch Engineer - Inference
Salary not disclosedRemote
- Engineering
- Bulgaria
- Full time
- Today
About the role
ElevenLabs is seeking a Research Engineer to join their research team, focusing on the deployment and optimization of frontier AI models in production. The role involves owning the systems that transition research breakthroughs into real-time products used by millions of users. You will be responsible for ensuring models are served fast, reliably, and at scale.
Responsibilities
- Deploy state-of-the-art models to production and own the path from research checkpoint to serving infrastructure.
- Optimize inference performance across the stack, including latency, throughput, and cost, using techniques such as quantization, distillation, KV-cache optimization, batching strategies, and custom kernels.
- Build and tune high-performance serving systems for real-time, streaming workloads.
- Create tooling and infrastructure that enables researchers to ship new models to production quickly, safely, and with confidence in their performance characteristics.
Required skills
- ML model deployment
- GPU programming
- Inference optimization
- CUDA
- Triton
- TensorRT
- vLLM
- SGLang
Qualifications
- Equivalent practical experience
Benefits
- Annual professional development stipend
- Annual social travel stipend
- Annual company offsite
- Monthly co-working stipend
About the Company
ElevenLabs is an AI research and product company that launched in January 2023 with the first human-like AI voice model. They serve millions of users and thousands of businesses, with platforms spanning voice agents, creative tools, and developer APIs.