Back to results

tensordyne

Forward Deployed Inference Engineer

Salary not disclosedOnsite

  • Engineering
  • Munich
  • Full time
  • 6d ago

About the role

The Forward Deployed Inference Engineer will bridge the gap between customer model requirements and Tensordyne's high-performance AI inference hardware. This role involves optimizing model performance, managing customer deployments, and collaborating with internal engineering teams to refine the platform based on real-world usage. You will own the technical path from initial workload assessment to successful production rollout.

Responsibilities

  • Turn customer workloads into fast, credible performance answers through profiling and benchmarking.
  • Define relevant KPIs, compare against competitive baselines, and keep evaluation methodology current.
  • Convert and bring up customer models on the Tensordyne stack, validate numerical quality, and identify performance bottlenecks.
  • Work directly with customers and partners on technical PoCs, integration, deployment, and debugging.
  • Translate requirements into measurable acceptance criteria for quality, latency, and throughput.
  • Track profiling-to-hardware accuracy by comparing results with actual hardware deployments and flagging missing capabilities.
  • Turn repeated customer-specific learnings into reusable tooling, documentation, and product improvements.

Required skills

  • AI models
  • Inference systems
  • LLMs
  • Python
  • PyTorch
  • Profiling
  • Benchmarking
  • Model optimization

Nice to have

  • vLLM
  • SGLang
  • Quantization
  • Distributed inference
  • Kubernetes
  • Rust

Qualifications

  • Strong hands-on experience with AI models and inference systems
  • Experience profiling, benchmarking, or optimizing model inference
  • Experience working directly with customers or external technical partners
  • Experience bringing models up on new accelerators or non-standard hardware

About the Company

Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world's most demanding generative AI workloads. The platform integrates purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture for large-scale AI inference.