gridscale GmbH
Visit websiteSenior Site Reliability Engineer - Ceph / On-Prem Cloud Storage
Salary not disclosedOnsite
- Engineering
- Germany
- Full time
- Yesterday
About the role
You will help build, operate, and industrialize the storage foundation of an on-premise cloud platform. As part of an experienced team, you will own Ceph clusters for block, object, and file storage services while influencing storage architecture, hardware lifecycle, and automation strategy. The role involves working across bare metal, networking, OpenStack, Kubernetes, and GitOps to ensure storage is predictable, durable, and highly automated.
Responsibilities
- Design, operate and evolve Ceph clusters for block, object and file storage.
- Own the storage hardware lifecycle end to end from qualification through maintenance and retirement.
- Automate the provisioning of storage nodes on OpenStack-managed bare metal.
- Run upgrades, rebalancing and reconfigurations as routine production operations.
- Own capacity planning and growth forecasting using telemetry and monitoring.
- Use LLMs and agentic tools to support development, testing, and automation.
Required skills
- Ceph
- SRE
- Storage Engineering
- Linux
- Bare metal
- Kubernetes
- OpenStack
- Ansible
- Terraform
- GitOps
- Python
- Go
Nice to have
- RGW
- S3
- RBD mirroring
- CephFS
- ceph-csi
- VLAN
- BGP
- NVMe
- BlueStore
Qualifications
- Several years of hands-on experience as an SRE, Storage Engineer or Platform Engineer running production storage infrastructure
Benefits
- 32 vacation days
- Employer-funded pension plan
- Insurance package
- Public transportation subsidy
- Annual sports contribution
- Corporate discounts
- Cargo bike leasing
- Company events
About the Company
gridscale GmbH is an international company that operates in a highly innovative environment with cutting-edge technologies. They emphasize team spirit across departments and national borders.