Nebius
Visit websiteSenior ML Engineer (Token Factory)
Salary not disclosedRemote
- Engineering
- Czechia
- Full time
- Yesterday
About the role
The Token Factory team at Nebius is building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. You will work on maximizing throughput, minimizing latency, and optimizing cost-per-token across tens of thousands of GPUs. This role involves identifying inference bottlenecks, implementing novel speculative decoding architectures, and designing low-precision training and inference pipelines.
Responsibilities
- Identify LLM inference bottlenecks to drive production speedups.
- Optimize performance for a wide range of LLM architectures at scale.
- Implement novel speculative decoding architectures and contribute to open-source inference engines.
- Design and productionize low-precision training and inference pipelines.
- Profile GPU workloads using industry-standard tools to ensure efficiency.
Required skills
- Machine Learning
- Transformer architecture
- GPU profiling
- PyTorch
- LLM
- MHA
- RoPE
- KV-cache
- Flash Attention
- Quantisation
- Python
- Deep learning frameworks
- CI/CD
- Unit testing
Nice to have
- vLLM
- SGLang
- TensorRT-LLM
- Triton
- Cute
- CUTLASS
- CUDA
- Distributed systems
Benefits
- Competitive compensation
- Career growth and learning opportunities
- Flexibility and ownership
- Collaborative and innovative culture
About the Company
Nebius is a full-stack AI cloud platform supporting developers and enterprises from data and model training through to production deployment. Listed on Nasdaq and headquartered in Amsterdam, the company operates a global footprint with R&D hubs across Europe, the UK, North America, and Israel.