Fact Finder
Visit websiteTeam Lead - Site Reliability Engineering
Salary not disclosedHybrid
- Engineering
- Berlin
- Full time
- 3d ago
About the role
As a Team Lead for Site Reliability Engineering, you will be responsible for the reliability, scalability, and cost-efficiency of hosting environments. You will lead the team through a transformation toward a Kubernetes-based platform on Harvester, managing the transition from on-premise to cloud-ready architectures. This role involves both technical leadership in platform engineering and the disciplinary management of a four-person team.
Responsibilities
- Manage the operational health of hosting environments across on-premise and cloud, including availability, performance, and incident management.
- Drive the modernization toward Kubernetes on Harvester, including cluster topology, storage, networking, and disaster recovery.
- Build a production-ready Kubernetes platform using GitOps, observability, and policy guardrails.
- Design the NG Search Operator and implement auto-scaling solutions.
- Define the on-premise hybrid model, ensuring architecture portability and cost-effective capacity planning.
- Lead and develop the hosting team, fostering a culture of technical standards and ownership.
- Integrate AI tools into operations for diagnosis, automation, and monitoring.
Required skills
- Kubernetes
- Platform Engineering
- Infrastructure Engineering
- GitOps
- Observability
- Auto-scaling
- Stakeholder Management
- Team Leadership
- AI tools
Nice to have
- Harvester
- HCI
- KubeVirt
- vSphere
- ESXi
- OpenStack
- Kubernetes Operators
Qualifications
- Background in Infrastructure or Platform Engineering
- Experience with migration from Bare Metal or VMs to Kubernetes
- Experience with on-prem hybrid architectures
Benefits
- Competitive salary
- Modern equipment
- Learning budget
- Team events
About the Company
FactFinder develops product discovery technology for eCommerce, used by leading online shops in Europe with products like Next Generation and Infinity.