Back to results

Senior ML Engineer (Token Factory)

Salary not disclosedRemote

  • Engineering
  • Czechia
  • Full time
  • Yesterday
Newly posted

About the role

The Token Factory team at Nebius is building a high-performance inference and fine-tuning platform designed to push foundation models to their hardware limits. You will work on maximizing throughput, minimizing latency, and optimizing cost-per-token across tens of thousands of GPUs. This role involves identifying inference bottlenecks, implementing novel speculative decoding architectures, and designing low-precision training and inference pipelines.

Responsibilities

  • Identify LLM inference bottlenecks to drive production speedups.
  • Optimize performance for a wide range of LLM architectures at scale.
  • Implement novel speculative decoding architectures and contribute to open-source inference engines.
  • Design and productionize low-precision training and inference pipelines.
  • Profile GPU workloads using industry-standard tools to ensure efficiency.

Required skills

  • Machine Learning
  • Transformer architecture
  • GPU profiling
  • PyTorch
  • LLM
  • MHA
  • RoPE
  • KV-cache
  • Flash Attention
  • Quantisation
  • Python
  • Deep learning frameworks
  • CI/CD
  • Unit testing

Nice to have

  • vLLM
  • SGLang
  • TensorRT-LLM
  • Triton
  • Cute
  • CUTLASS
  • CUDA
  • Distributed systems

Benefits

  • Competitive compensation
  • Career growth and learning opportunities
  • Flexibility and ownership
  • Collaborative and innovative culture

About the Company

Nebius is a full-stack AI cloud platform supporting developers and enterprises from data and model training through to production deployment. Listed on Nasdaq and headquartered in Amsterdam, the company operates a global footprint with R&D hubs across Europe, the UK, North America, and Israel.