Kimchi
Senior ML Engineer
Salary not disclosed5+ yearsRemote
- Engineering
- United Kingdom
- Full time
- Yesterday
About the role
This role focuses on optimizing LLM inference performance, specifically targeting throughput, latency, and KV cache utilization. You will lead the technical direction for inference optimization, building the layer between models and hardware to ensure cost-efficient and high-performing production environments.
Responsibilities
- Push throughput by implementing continuous batching, speculative decoding, and kernel-level tuning.
- Cut latency by profiling and identifying bottlenecks in compute, memory, and networking.
- Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
- Perform quantization across weights, activations, and KV while ensuring quality maintenance.
- Shrink cold starts and memory footprints to improve scalability.
- Scale inference across nodes using distributed topologies and network-aware placement.
- Set the technical direction for benchmarking and infrastructure development.
Required skills
- Python
- vLLM
- SGLang
- TensorRT-LLM
- PyTorch
- Kubernetes
- Distributed systems
- Quantization
Qualifications
- 5+ years building real ML systems with depth in inference or training infrastructure
Benefits
- Competitive salary
- Equity options
- Learning budget
- Annual hackathon
- Team-building budget
- Equipment budget
- Extra days off
About the Company
Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale by embedding autonomous decision-making into Kubernetes environments. The company serves over 2,100 customers and recently achieved unicorn status.