Contabo
Visit websitePlatform & Reliability Engineer
Salary not disclosed7+ yearsHybrid
- Engineering
- Germany
- Full time
- Yesterday
About the role
As a Platform & Reliability Engineer, you will take architectural ownership of shared infrastructure services including API gateways, persistent storage, secrets management, and observability. You will be responsible for identifying and remediating weak points in existing systems, establishing formal on-call processes, and maturing observability practices. This role serves as a technical authority across multiple development teams to ensure infrastructure stability and knowledge distribution.
Responsibilities
- Take architectural ownership of shared infrastructure services including API gateway, ingress, persistent storage, and secrets management.
- Identify and remediate weak points in existing, partially documented systems.
- Design and establish an on-call process and incident runbooks.
- Mature observability practices by rolling out distributed tracing, SLOs/SLIs, and dashboards.
- Advise across teams on shared infrastructure services and document technical knowledge.
Required skills
- SRE
- Platform Engineering
- Infrastructure
- Ceph
- Kubernetes
- Longhorn
- Kong
- Nginx Ingress
- Cloudflare
- WAF
- System design
- Load balancing
- Caching
- Sharding
- Replication
Nice to have
- Prometheus
- Grafana
- OpenTelemetry
- Vault
- Keycloak
- NATS
- Proxmox
- OpenStack
- German language
Qualifications
- 7+ years of experience in platform, infrastructure, or SRE roles
- Hands-on production experience with distributed storage systems
- Professional fluency in English
Certifications
- CKA
- CKS
Benefits
- Flexible hours
- Workation across the EU
- EGYM Wellpass
- Corporate benefits program
- 30 days of vacation
- Volunteer Day
About the Company
Contabo is an international company that values trust, direct communication, and an open feedback culture. They focus on providing space for new ideas, ownership, and impact in an innovative tech environment.