Lumaai
Visit websiteResearch Scientist / Engineer – Performance Optimization
Salary not disclosedOnsite
- Engineering
- London
- Full time
- Today
About the role
You will be responsible for making Luma's multimodal models fast by profiling and optimizing GPU, CPU, and accelerator code. This role focuses on deep performance work, including writing kernels and operations to ensure efficient training and scalable deployment without sacrificing quality.
Responsibilities
- Profile and optimize GPU/CPU/accelerator code for maximum utilization and minimal latency.
- Write high-performance PyTorch, Triton, and CUDA, dropping to custom operations when needed.
- Develop fused kernels and leverage tensor cores and modern hardware features across platforms.
- Optimize model architectures and implementations for distributed multi-node production deployment.
- Build performance monitoring and analysis tools and automation.
- Research and implement cutting-edge optimization techniques for transformer models.
Required skills
- Triton
- CUDA
- GPU optimization
- PyTorch
- Profiling tools
- NVIDIA Nsight
- Transformer architectures
- Attention mechanisms
Nice to have
- torch.compile
- TensorRT
- ONNX
- XLA
- Inference optimization
- Kernel fusion
- Warp-level intrinsics
About the Company
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence.