Back to results

Research Scientist / Engineer – Training Infrastructure

Salary not disclosedOnsite

  • Engineering
  • London
  • Full time
  • Today
Newly posted

About the role

You will build and optimize distributed systems for training large-scale multimodal models across thousands of GPUs. This role focuses on developing reliable, efficient, and scalable infrastructure to support research innovation. You will work on advanced parallelism, training stability, and cluster utilization.

Responsibilities

  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization techniques including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel.
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.

Required skills

  • PyTorch
  • CUDA
  • Distributed systems
  • FSDP
  • GPU clusters
  • Networking
  • Storage systems
  • NCCL
  • MPI

Nice to have

  • Linux systems administration
  • Scripting
  • Containerization
  • Orchestration
  • Cloud infrastructure

About the Company

Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence.

Research Scientist / Engineer – Training Infrastructure at Lumaai · Grasshire