Mistral.ai
Research Engineer - Eval Platform
Salary not disclosed4+ yearsOnsite
- Engineering
- Paris
- Full time
- Today
About the role
As a Research Engineer on the Eval Platform team, you will build the infrastructure used to measure model quality, ensuring it is reliable, reproducible, and fast. You will work closely with researchers to turn new evaluation needs into robust, shared tooling that supports the development of frontier AI models.
Responsibilities
- Build systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve.
- Run evaluations at scale across our GPU clusters, from model serving to scoring.
- Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day.
- Catch broken or noisy evals before they mislead research decisions.
- Support evaluation of agentic, multi-turn and tool-using models.
- Work closely with researchers to turn new evaluation needs into robust, shared tooling.
Required skills
- Python
- Software design
- CI/CD
- GPU clusters
- Slurm
- Kubernetes
- Ray
- LLM inference
Nice to have
- Evaluation harnesses
- Benchmarks
- vLLM
- SGLang
- Agentic environments
- RL environments
- Statistics
- Variance estimation
- Significance testing
Qualifications
- Master's or PhD in Computer Science
- Equivalent practical experience
Benefits
- Healthcare coverage
- Parental leave
- Retirement plans
- Relocation support
- Wellness programs
- Meal allowances
- Transportation allowances
About the Company
Mistral provides full-stack AI solutions, from frontier models to developer tools, applications, and compute. The company partners with enterprises across high-stakes industries to co-create customized AI systems.