Back to results

Mistral.ai

Research Engineer, Inference Foundation

Salary not disclosedOnsite

  • Engineering
  • Paris
  • Full time
  • Yesterday
Newly posted

About the role

The Inference Foundation team owns the core of Mistral's inference stack, including the inference engine, orchestration, and release machinery. This role focuses on optimizing the inference stack for throughput and latency, ensuring elastic capacity, and powering the serving infrastructure for frontier model training. You will work on production LLM serving, engine development, and capacity engineering to maintain high-performance, production-grade systems.

Responsibilities

  • Develop and fix the core of the inference stack including engine and orchestrator configuration and tuning.
  • Own the release process for the serving stack through automated performance gates and progressive rollout.
  • Drive improvements and fixes upstream in open-source engines.
  • Optimize serving efficiency across the fleet to reduce startup times and improve caching.
  • Optimize and maintain serving topology including connectivity and routing.
  • Build serving infrastructure to power RL and post-training for frontier models.

Required skills

  • ML/LLM services
  • vLLM
  • SGLang
  • TensorRT-LLM
  • Inference internals
  • Distributed serving architectures
  • CUDA
  • NCCL
  • Python
  • PyTorch
  • Kubernetes
  • InfiniBand
  • RDMA

Nice to have

  • MoE models
  • Triton
  • Nsight Systems
  • Nsight Compute
  • Rust
  • C++

Benefits

  • Healthcare coverage
  • Parental leave
  • Retirement plans
  • Relocation support
  • Wellness programs
  • Meal allowance
  • Transportation allowance

About the Company

Mistral provides full-stack AI solutions ranging from frontier models to developer tools, applications, and compute. The company partners with enterprises across high-stakes industries to co-create customized AI systems.

Research Engineer, Inference Foundation at Mistral.ai · Grasshire