Jobgether
Visit websiteAI Research Engineer (Kernel & Inference Optimization)
Salary not disclosedOnsite
- Engineering
- Switzerland
- Full time
- 3d ago
Newly posted
About the role
This role sits at the intersection of AI research, systems engineering, and high-performance model inference. You will develop and optimize model-serving architectures for advanced AI systems, focusing on latency, throughput, and memory efficiency across diverse hardware environments including mobile and edge devices. The position involves hands-on research, custom GPU kernel development, and iterative optimization to translate research into measurable performance improvements.
Responsibilities
- Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
- Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
- Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
- Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
- Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
- Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
- Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
- Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
Required skills
- Metal Shading Language
- Kernel optimization
- Inference optimization
- GPU kernels
- Model-serving architectures
- Diffusion models
- Vision Transformers
- Pruning
- Quantization
- Flash Attention
- KV caching
- Speculative decoding
- Tensor parallelism
- Pipeline parallelism
Qualifications
- Degree in Computer Science or a related technical field
- PhD in NLP, Machine Learning, or a related discipline
Benefits
- Remote-first working environment
- Exposure to cutting-edge AI research
- Collaborative research-driven environment