Lumaai
Visit websiteResearch Scientist / Engineer – Training Infrastructure
Salary not disclosedOnsite
- Engineering
- London
- Full time
- Today
About the role
You will build and optimize distributed systems for training large-scale multimodal models across thousands of GPUs. This role focuses on developing reliable, efficient, and scalable infrastructure to support research innovation. You will work on advanced parallelism, training stability, and cluster utilization.
Responsibilities
- Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs.
- Research and implement advanced parallelization techniques including FSDP, Tensor Parallel, Pipeline Parallel, and Expert Parallel.
- Build monitoring, visualization, and debugging tools for large-scale training runs.
- Optimize training stability, convergence, and resource utilization across massive clusters.
Required skills
- PyTorch
- CUDA
- Distributed systems
- FSDP
- GPU clusters
- Networking
- Storage systems
- NCCL
- MPI
Nice to have
- Linux systems administration
- Scripting
- Containerization
- Orchestration
- Cloud infrastructure
About the Company
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence.