tensordyne
Forward Deployed Inference Engineer
Salary not disclosedOnsite
- Engineering
- Munich
- Full time
- 4d ago
About the role
The Forward Deployed Inference Engineer will bridge the gap between customer needs and Tensordyne's AI inference platform. This role involves optimizing generative AI models for high-performance hardware, managing technical customer deployments, and collaborating with internal engineering teams to improve system capabilities.
Responsibilities
- Turn customer workloads into fast, credible performance answers through profiling and benchmarking.
- Convert and bring up customer models on the Tensordyne stack, validate numerical quality, and identify performance bottlenecks.
- Work directly with customers and partners on technical PoCs, integration, deployment, and debugging.
- Track profiling-to-hardware accuracy by comparing simulation results with actual hardware deployments.
- Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements.
Required skills
- AI models
- Inference systems
- LLMs
- Python
- PyTorch
- Profiling
- Benchmarking
- Model optimization
Nice to have
- vLLM
- SGLang
- Customer-facing
- Hardware acceleration
- Quantization
- Distributed inference
- Kubernetes
- Rust
About the Company
Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world’s most demanding generative AI workloads. Their platform integrates purpose-built silicon, new AI math, and optimized networking for large-scale AI inference.