Nebius
Visit websiteSenior Machine Learning Engineer, LLM Inference Optimization
Salary not disclosedOnsite
- Engineering
- Switzerland
- Full time
- 5d ago
About the role
Nebius is seeking a Senior Machine Learning Engineer to join the Applied AI team and optimize inference services for frontier models. You will own model and endpoint optimization from artifacts through production, focusing on latency, throughput, memory efficiency, and cost per token. This hands-on role involves diagnosing complex serving problems and delivering measurable performance improvements in production environments.
Responsibilities
- Own optimization work for specific model families, customer endpoints, or serving backends.
- Run engine comparisons and recommend practical serving configurations for specific workloads.
- Debug model quality or performance regressions during production rollouts.
- Optimize LLM and VLM endpoints for latency, throughput, memory efficiency, GPU utilization, quality, and cost per token.
- Deploy, configure, benchmark, and extend inference engines such as vLLM, SGLang, TensorRT-LLM, Triton Inference Server, or similar systems.
- Build and productionize model-compression workflows, including quantization, quantization-aware training, distillation, low-bit serving, and accuracy recovery.
- Implement or integrate speculative decoding, draft-model approaches, KV-cache optimization, prefix caching, chunked prefill, continuous batching, and disaggregated prefill/decode serving.
- Build reproducible benchmark harnesses for TTFT, TPOT, tokens per second per GPU, p95/p99 latency, GPU memory, reliability, and cost per token.
- Partner with GPU kernel engineers and platform engineers to diagnose bottlenecks across model code, kernels, runtime, scheduler, gateway, and cluster layers.
- Write clear design docs, performance reports, rollout plans, and customer-facing technical explanations.
Required skills
- Python
- PyTorch
- LLM
- VLM
- Transformer inference
- vLLM
- SGLang
- TensorRT-LLM
- Triton Inference Server
- Ray Serve
- KServe
- KV cache
- Attention mechanisms
- Memory bandwidth
- Parallelism
Nice to have
- Quantization-aware training
- Post-training quantization
- FP8
- INT8
- INT4
- NVFP4
- MXFP4
- AWQ
- GPTQ
- SmoothQuant
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
- Opportunity to work on impactful AI projects
- International environment
About the Company
Nebius is a cloud infrastructure company building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment. Headquartered in Amsterdam and listed on Nasdaq, the company has a global footprint with R&D hubs across Europe, the UK, North America, and Israel.