tensordyne
Forward Deployed Inference Engineer
Salary not disclosedOnsite
- Engineering
- Munich
- Full time
- 6d ago
About the role
The Forward Deployed Inference Engineer will bridge the gap between customer model requirements and Tensordyne's high-performance AI inference hardware. This role involves optimizing model performance, managing customer deployments, and collaborating with internal engineering teams to refine the platform based on real-world usage. You will own the technical path from initial workload assessment to successful production rollout.
Responsibilities
- Turn customer workloads into fast, credible performance answers through profiling and benchmarking.
- Define relevant KPIs, compare against competitive baselines, and keep evaluation methodology current.
- Convert and bring up customer models on the Tensordyne stack, validate numerical quality, and identify performance bottlenecks.
- Work directly with customers and partners on technical PoCs, integration, deployment, and debugging.
- Translate requirements into measurable acceptance criteria for quality, latency, and throughput.
- Track profiling-to-hardware accuracy by comparing results with actual hardware deployments and flagging missing capabilities.
- Turn repeated customer-specific learnings into reusable tooling, documentation, and product improvements.
Required skills
- AI models
- Inference systems
- LLMs
- Python
- PyTorch
- Profiling
- Benchmarking
- Model optimization
Nice to have
- vLLM
- SGLang
- Quantization
- Distributed inference
- Kubernetes
- Rust
Qualifications
- Strong hands-on experience with AI models and inference systems
- Experience profiling, benchmarking, or optimizing model inference
- Experience working directly with customers or external technical partners
- Experience bringing models up on new accelerators or non-standard hardware
About the Company
Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world's most demanding generative AI workloads. The platform integrates purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture for large-scale AI inference.