Back to results

Kimchi

Senior ML Engineer

Salary not disclosed5+ yearsRemote

  • Engineering
  • Full time
  • Yesterday
Newly posted

About the role

This role focuses on optimizing LLM inference performance, specifically targeting throughput, latency, and KV cache utilization. You will lead the technical direction for inference optimization, building systems that automatically match workloads to the most efficient model and serving configurations. The position requires deep expertise in both model architecture and hardware interaction to improve customer performance and company margins.

Responsibilities

  • Push throughput by implementing continuous batching, speculative decoding, and kernel-level tuning.
  • Cut latency by profiling and identifying bottlenecks in compute, memory, and networking.
  • Optimize KV cache usage through paged attention, prefix caching, and eviction policies.
  • Perform quantization across weights, activations, and KV while ensuring quality standards.
  • Shrink cold starts and memory footprints to improve scalability.
  • Scale inference across nodes using distributed topologies and network-aware placement.
  • Set the technical direction for benchmarking and infrastructure development.

Required skills

  • Python
  • vLLM
  • SGLang
  • TensorRT-LLM
  • PyTorch
  • CUDA
  • Kubernetes
  • Distributed systems
  • Quantization
  • Inference optimization

Benefits

  • Competitive salary
  • Equity options
  • Learning budget
  • Annual hackathon
  • Team-building budget
  • Equipment budget
  • Extra days off

About the Company

Cast AI is an automation platform that operates cloud-native and AI infrastructure at scale by embedding autonomous decision-making into Kubernetes environments. The company serves over 2,100 customers and recently achieved unicorn status.

Senior ML Engineer at Kimchi · Grasshire