Back to results

Mistral.ai

Research Engineer - Eval Platform

Salary not disclosed4+ yearsOnsite

  • Engineering
  • Paris
  • Full time
  • Today
Newly posted

About the role

As a Research Engineer on the Eval Platform team, you will build the infrastructure used to measure model quality, ensuring it is reliable, reproducible, and fast. You will work closely with researchers to turn new evaluation needs into robust, shared tooling that supports the development of frontier AI models.

Responsibilities

  • Build systems that keep eval results reproducible and comparable over time, as models, benchmarks and code evolve.
  • Run evaluations at scale across our GPU clusters, from model serving to scoring.
  • Make eval results easy to access, explore and trust, through APIs and dashboards that researchers use every day.
  • Catch broken or noisy evals before they mislead research decisions.
  • Support evaluation of agentic, multi-turn and tool-using models.
  • Work closely with researchers to turn new evaluation needs into robust, shared tooling.

Required skills

  • Python
  • Software design
  • CI/CD
  • GPU clusters
  • Slurm
  • Kubernetes
  • Ray
  • LLM inference

Nice to have

  • Evaluation harnesses
  • Benchmarks
  • vLLM
  • SGLang
  • Agentic environments
  • RL environments
  • Statistics
  • Variance estimation
  • Significance testing

Qualifications

  • Master's or PhD in Computer Science
  • Equivalent practical experience

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowances
  • Transportation allowances

About the Company

Mistral provides full-stack AI solutions, from frontier models to developer tools, applications, and compute. The company partners with enterprises across high-stakes industries to co-create customized AI systems.