Back to results

Senior Software Engineer, Chaos Engineering

Salary not disclosedHybrid

  • Engineering
  • Paris
  • Full time
  • 2d ago
Newly posted

About the role

The Chaos Engineering team builds systems to identify and address reliability weaknesses before they cause outages. As a Senior Software Engineer, you will focus on zonal resilience, fault injection, and gameday orchestration while collaborating with engineering teams to design systems that safely exercise production failure modes. You will also leverage AI and automation to improve resilience testing and remediation processes.

Responsibilities

  • Build zonal-resilience automation that coordinates safe workload evacuations, switchovers, and recovery in partnership with the teams that own affected services.
  • Design and build fault-injection systems for production environments, including infrastructure- and application-level testing, incident replay, and controlled resilience experiments.
  • Develop safeguards such as blast-radius controls, kill switches, validation mechanisms, and rollback paths that keep production experiments contained and reversible.
  • Build agents and automation that help propose failure scenarios, triage experiment results, and connect reliability findings to tracked remediation and verification.
  • Lead gamedays from hypothesis and scenario design through execution, documented findings, remediation tracking, and validation of completed fixes.
  • Design and implement reliable distributed systems, including gRPC services, Kubernetes controllers, and shared platform components, while contributing to technical design and mentoring other engineers.

Required skills

  • Distributed systems
  • Kubernetes
  • gRPC
  • Fault injection
  • Reliability engineering

Nice to have

  • Chaos engineering
  • Zonal failover
  • AI-assisted operational workflows
  • Traffic interception
  • Observability systems

About the Company

Datadog is the leading observability and security platform for the AI era, providing businesses with unified visibility across the technology stack to manage complexity at scale. It brings applications, infrastructure, data, models, and security into one place, using AI to detect and resolve issues before they impact customers.

Senior Software Engineer, Chaos Engineering at Datadog · Grasshire