Jobgether
Visit websiteAI Research Engineer (Kernel & Inference Optimization)
Salary not disclosedOnsite
- Engineering
- Switzerland
- Full time
- 3d ago
Newly posted
About the role
This role sits at the intersection of AI research, systems engineering, and high-performance model inference. You will be responsible for developing and optimizing model-serving architectures for advanced AI systems across various hardware environments, including mobile and edge devices. The position involves hands-on research and low-level engineering to improve latency, throughput, and memory efficiency for complex architectures like diffusion models and vision transformers.
Responsibilities
- Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
- Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
- Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
- Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
- Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
- Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
- Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
- Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.
Required skills
- Metal Shading Language
- GPU kernel optimization
- Inference optimization
- Model-serving architectures
- Distributed inference
- Diffusion models
- Vision Transformers
- Benchmarking
- C++
- Python
Qualifications
- Degree in Computer Science or a related technical field
- PhD in NLP, Machine Learning, or a related discipline
Benefits
- Remote-first working environment
- Exposure to cutting-edge AI research
- Work with advanced model architectures