Back to results

Senior Applied Scientist, Efficient LLM Inference & Model Optimization

Salary not disclosedOnsite

  • Engineering
  • Switzerland
  • Full time
  • 3d ago
Newly posted

About the role

Nebius is seeking a Senior Applied Scientist to join the Token Factory team to address frontier inference bottlenecks. The role involves designing rigorous experiments, writing high-quality code, and collaborating with engineers to transition research into production-ready inference capabilities.

Responsibilities

  • Own focused research projects from hypothesis through experiment, ablation, prototype, and production handoff.
  • Prepare internal reports, technical blogs, or papers when the work is externally credible.
  • Partner directly with MLEs to ensure research prototypes become usable production components.
  • Define and execute research programs in efficient LLM and VLM inference with measurable production impact.
  • Invent, evaluate, and productionize methods for quantization, QAT, distillation, speculative decoding, KV-cache reuse, KV-cache compression, long-context inference, MoE routing, and model/runtime co-optimization.
  • Build high-quality prototypes in PyTorch, Triton, CUDA-adjacent tooling, or inference-serving frameworks.
  • Design rigorous evaluation methodology covering quality, latency, throughput, numerical stability, memory footprint, tail latency, and cost per token.
  • Mentor engineers and scientists on experimental design, scientific rigor, and model/system tradeoffs.

Required skills

  • Machine Learning
  • LLM
  • VLM
  • Transformer inference
  • Model compression
  • Quantization
  • Distillation
  • Python
  • PyTorch
  • Experimental design

Nice to have

  • vLLM
  • SGLang
  • TensorRT-LLM
  • NVIDIA Dynamo
  • FlashAttention
  • FlashInfer
  • Triton
  • CUDA
  • SFT
  • DPO

Qualifications

  • PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field.

Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture

About the Company

Nebius is building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment. Listed on Nasdaq, the company has a global footprint with R&D hubs across Europe, the UK, North America and Israel.

Senior Applied Scientist, Efficient LLM Inference & Model Optimization at Nebius · Grasshire