
On-site
Senior Site Reliability Engineer
Ireland, Ireland
full-timeOn-sitePosted 2 Sept 2026
โ All rolesKubernetesPythonObservabilitySystem Design
We are hiring a Senior Site Reliability Engineer (SRE) to champion system availability, operational resilience, and observability across production workloads.
Tech stack: Kubernetes, Python, Prometheus, Grafana, Terraform, AWS, Chaos Engineering, Incident Response
What you will work on:
- Define, track, and maintain SLOs, SLIs, and Error Budgets across core production services.
- Build modern observability stacks utilizing Prometheus, Grafana, OpenTelemetry, and structured logging tools.
- Lead major incident management, post-mortem blameless reviews, and root-cause prevention efforts.
- Develop custom automation scripts and operator utilities in Python or Go to eliminate system toil.
What we are looking for:
- Solid track record in SRE or DevOps roles handling mission-critical cloud deployments.
- In-depth expertise in Kubernetes operations, Linux performance debugging, network analysis, and observability.
- Strong incident management background paired with excellent automation and scripting capabilities.