Back to results

Kimchi

Senior ML Engineer

Salary not disclosed5+ yearsRemote

  • Engineering
  • United Kingdom
  • Full time
  • Yesterday
Newly posted

About the role

This role focuses on optimizing LLM inference performance, specifically targeting throughput, latency, and KV cache utilization. You will lead the technical direction for inference optimization, building the layer between models and hardware to ensure cost-efficient and high-performing production environments.

Responsibilities

  • Push throughput by implementing continuous batching, speculative decoding, and kernel-level tuning.
  • Cut latency by profiling and identifying bottlenecks in compute, memory, and networking.
  • Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
  • Perform quantization across weights, activations, and KV while ensuring quality maintenance.
  • Shrink cold starts and memory footprints to improve scalability.
  • Scale inference across nodes using distributed topologies and network-aware placement.
  • Set the technical direction for benchmarking and infrastructure development.

Required skills

  • Python
  • vLLM
  • SGLang
  • TensorRT-LLM
  • PyTorch
  • Kubernetes
  • Distributed systems
  • Quantization

Qualifications

  • 5+ years building real ML systems with depth in inference or training infrastructure

Benefits

  • Competitive salary
  • Equity options
  • Learning budget
  • Annual hackathon
  • Team-building budget
  • Equipment budget
  • Extra days off

About the Company

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale by embedding autonomous decision-making into Kubernetes environments. The company serves over 2,100 customers and recently achieved unicorn status.

Senior ML Engineer at Kimchi · Grasshire