Back to results

Software Engineer, Inference

Salary not disclosedOnsite

  • Engineering
  • London
  • Full time
  • Today
Newly posted

About the role

This role focuses on large-scale inference systems, including scheduling, fleet management, and deployment pipelines. You will be responsible for integrating new model architectures into the inference engine and ensuring the reliability and efficiency of GPU fleets across clusters.

Responsibilities

  • Ship new model architectures by integrating them into the inference engine.
  • Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
  • Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
  • Automate, test, and maintain inference services for maximum uptime and reliability.
  • Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
  • Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.

Required skills

  • Python
  • System architecture
  • PyTorch
  • Hugging Face
  • vLLM
  • SGLang
  • TensorRT-LLM
  • Queues
  • Scheduling
  • Traffic control
  • Fleet management
  • Linux
  • Docker
  • Kubernetes
  • Redis

Nice to have

  • RDMA
  • RoCE
  • InfiniBand
  • NVLink
  • CUDA
  • FFmpeg
  • Multimedia processing

About the Company

Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence.