Senior Platform Engineer AI & Observability
Salary not disclosedOnsite
- Germany
- Full time
- Yesterday
About the role
The Senior Platform Engineer will design, deploy, and operate scalable AI, observability, and cloud-native platforms using Kubernetes and HPC technologies. This role involves building AI services, maintaining monitoring and security solutions, and providing technical leadership to support federated research infrastructures. The position is based at the Leibniz Supercomputing Centre, contributing to advanced IT services for research.
Responsibilities
- Design, deploy, and operate scalable AI, observability, and cloud-native platforms based on Kubernetes and HPC technologies.
- Build and optimize AI services, LLM inference platforms, and GPU-enabled workloads.
- Develop and maintain monitoring, logging, tracing, and security solutions using open-source technologies.
- Create standardized deployment workflows, automation, and platform best practices.
- Enable reliable, secure, and multi-tenant operation of federated research infrastructures.
- Collaborate with project partners and provide technical leadership in architecture, implementation, and operations.
Required skills
- Linux
- Docker
- Kubernetes
- Cloud-native technologies
- Observability
- Monitoring
- Logging
- Tracing
- Security concepts
- Python
- Go
- Bash
- DevOps
- MLOps
- Infrastructure automation
Nice to have
- AI
- Machine learning
- LLMs
- vLLM
- Triton
- Ollama
- Llama.cpp
- Prometheus
- Grafana
- Helm
Qualifications
- Master’s degree in Computer Science, Data Science, Computer Engineering, or a related field
Benefits
- Flexible working model
- Pension plan
- State-of-the-art work equipment
- Discounted sports facility membership
- Free parking
About the Company
LRZ is an institute of the Bavarian Academy of Science and Humanities that advances IT services for ground-breaking research. It provides a dynamic, cooperative, and innovative work environment with an international team of experts.