Kimchi
Senior ML Engineer
Salary not disclosed5+ yearsRemote
- Engineering
- Full time
- Yesterday
About the role
This role focuses on optimizing LLM inference performance, specifically targeting throughput, latency, and KV cache utilization. You will lead the technical direction for inference optimization, building systems that automatically match workloads to the most efficient model and serving configurations. The position requires deep expertise in both model architecture and hardware interaction to improve customer performance and company margins.
Responsibilities
- Push throughput by implementing continuous batching, speculative decoding, and kernel-level tuning.
- Cut latency by profiling and identifying bottlenecks in compute, memory, and networking.
- Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
- Perform quantization across weights, activations, and KV while ensuring quality standards.
- Shrink cold starts and memory footprints to improve scalability.
- Scale inference across nodes using distributed topologies and network-aware placement.
- Set the technical direction for benchmarking and infrastructure development.
Required skills
- Python
- vLLM
- SGLang
- TensorRT-LLM
- PyTorch
- CUDA
- Kubernetes
- Distributed systems
- Quantization
- Inference optimization
Benefits
- Competitive salary
- Equity options
- Learning budget
- Annual hackathon
- Team-building budget
- Equipment budget
- Extra days off
About the Company
Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale by embedding autonomous decision-making into Kubernetes environments. The company serves over 2,100 customers and recently achieved unicorn status.