Lumaai
Visit websiteSoftware Engineer, Inference
Salary not disclosedOnsite
- Engineering
- London
- Full time
- Today
About the role
This role focuses on large-scale inference systems, including scheduling, fleet management, and deployment pipelines. You will be responsible for integrating new model architectures into the inference engine and ensuring the reliability and efficiency of GPU fleets across clusters.
Responsibilities
- Ship new model architectures by integrating them into the inference engine.
- Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
- Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
- Automate, test, and maintain inference services for maximum uptime and reliability.
- Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
- Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.
Required skills
- Python
- System architecture
- PyTorch
- Hugging Face
- vLLM
- SGLang
- TensorRT-LLM
- Queues
- Scheduling
- Traffic control
- Fleet management
- Linux
- Docker
- Kubernetes
- Redis
Nice to have
- RDMA
- RoCE
- InfiniBand
- NVLink
- CUDA
- FFmpeg
- Multimedia processing
About the Company
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence.