Back to results

Lightningai

Senior Infrastructure Operations Engineer

£116,000 – £128,000Onsite

  • Engineering
  • London
  • Full time
  • Today
Newly posted

About the role

Lightning AI is seeking an experienced Senior Infrastructure Operations Engineer to help operate and scale the infrastructure behind one of the world's largest GPU fleets. The Infrastructure Operations team sits at the center of production reliability, working across Linux, bare-metal servers, GPUs, networking, storage, and cluster orchestration.

Responsibilities

  • Operate and troubleshoot large-scale GPU and bare-metal infrastructure across Linux, compute, networking, storage, and cluster orchestration.
  • Serve as a technical escalation point for complex infrastructure issues, driving problems from initial investigation through resolution and root cause analysis.
  • Own operational workflows across the infrastructure lifecycle, including provisioning, configuration, validation, maintenance, remediation, and decommissioning.
  • Build automation and internal tooling that eliminates repetitive operational work and enables the infrastructure fleet to scale efficiently.
  • Identify recurring failure modes and partner with Infrastructure Engineering to build more reliable, repeatable, and automated systems.
  • Improve provisioning, monitoring, diagnostics, and operational processes across a rapidly growing infrastructure footprint.
  • Partner closely with Infrastructure Engineering, Network Engineering, Data Center Operations, Customer Experience, and Platform Engineering to resolve issues that cross team or system boundaries.
  • Participate in a distributed primary/secondary on-call rotation supporting production infrastructure.

Required skills

  • Linux
  • Python
  • Go
  • Bash
  • Ansible
  • Kubernetes
  • Slurm
  • Networking
  • Observability

Nice to have

  • Bare-metal server infrastructure
  • GPU
  • HPC
  • NVIDIA GPUs
  • DCGM
  • InfiniBand
  • RoCE/RDMA
  • NVLink
  • PXE
  • BMC

Qualifications

  • Strong experience operating and troubleshooting Linux-based production infrastructure at scale
  • Experience troubleshooting complex infrastructure issues across compute, networking, storage, and operating systems
  • Strong systems and networking fundamentals

Benefits

  • Medical, dental, and vision coverage
  • Equity
  • Pension contributions
  • Unlimited PTO
  • Company-wide winter break
  • Paid parental and family leave
  • Professional development allowance
  • Wellness and work-from-home stipends
  • Sabbatical program
  • In-office meals

About the Company

Lightning AI is the company behind PyTorch Lightning, building an end-to-end platform for developing, training, and deploying AI systems. Through a merger with Voltage Park, the company combines developer-first software with cost-efficient, large-scale compute.

Senior Infrastructure Operations Engineer at Lightningai · Grasshire