Senior ML Engineer
KimchiThis role focuses on optimizing LLM inference performance, specifically targeting throughput, latency, and KV cache utilization. You will lead the technical direction for inference…
- Python
- vLLM
- SGLang
- TensorRT-LLM
- PyTorch
- Kubernetes
- +3 more