
Hybrid
Site Reliability Engineer
United Kingdom, United Kingdom
full-timeHybridPosted 2 Sept 2026
โ All rolesKubernetesPythonObservability
We are looking for a Site Reliability Engineer (SRE) to ensure high operational availability, system resilience, and observability across our platform.
Tech stack: Kubernetes, Python, Prometheus, Grafana, Terraform, AWS, Docker, Linux
What you will work on:
- Implement and manage platform observability tooling using Prometheus, Grafana, and open telemetry collectors.
- Configure alerting rules, operational dashboards, and incident triage runbooks.
- Participate in on-call rotation schedules, incident mitigation, and blameless post-mortem reviews.
- Automate system recovery tasks, chaos testing experiments, and infrastructure maintenance workflows.
What we are looking for:
- Practical experience in SRE, DevOps, or System Operations handling cloud applications.
- Hands-on technical knowledge of Kubernetes operations, Linux systems, Prometheus monitoring, and Python automation.
- Strong diagnostic skills and systematic troubleshooting methodology under operational pressure.