Back to results

Jobgether

Visit website

AI Research Engineer (Kernel & Inference Optimization)

Salary not disclosedOnsite

  • Engineering
  • Switzerland
  • Full time
  • 3d ago
Newly posted

About the role

This role sits at the intersection of AI research, systems engineering, and high-performance model inference. You will be responsible for developing and optimizing model-serving architectures for advanced AI systems across various hardware environments, including mobile and edge devices. The position involves hands-on research and low-level engineering to improve latency, throughput, and memory efficiency for complex architectures like diffusion models and vision transformers.

Responsibilities

  • Design and deploy advanced model-serving architectures optimized for high throughput, low latency, and efficient memory utilization.
  • Develop inference pipelines capable of operating effectively across diverse environments, including resource-constrained mobile devices and edge platforms.
  • Establish clear performance targets covering response latency, token generation speed, throughput, memory footprint, and reliability.
  • Build and execute controlled inference benchmarks in simulated and production environments, tracking latency, throughput, memory consumption, and error rates.
  • Create and maintain representative datasets and simulation scenarios for evaluating model performance under real-world and resource-constrained conditions.
  • Identify computational and memory bottlenecks across inference pipelines and implement solutions involving batching, networking, memory management, and other system-level optimizations.
  • Develop custom GPU kernels and compute shaders for mobile hardware, including solutions written in Metal Shading Language (MSL).
  • Apply advanced inference optimization techniques such as pruning, quantization, Flash Attention, KV caching, and speculative decoding.

Required skills

  • Metal Shading Language
  • GPU kernel optimization
  • Inference optimization
  • Model-serving architectures
  • Distributed inference
  • Diffusion models
  • Vision Transformers
  • Benchmarking
  • C++
  • Python

Qualifications

  • Degree in Computer Science or a related technical field
  • PhD in NLP, Machine Learning, or a related discipline

Benefits

  • Remote-first working environment
  • Exposure to cutting-edge AI research
  • Work with advanced model architectures
AI Research Engineer (Kernel & Inference Optimization) at Jobgether · Grasshire