About this role
Salary: £65,000 - 65,000 per year
Requirements: Experience in a Site Reliability Engineering, Production Engineering, Cloud Operations, or NOC environmentExposure to Linux systems administrationExposure to AWS cloud infrastructureExposure to Kubernetes and DockerExperience with production support and incident managementScripting experience in Python, Bash, or GoFamiliarity with monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk, or CloudWatchUnderstanding of networking fundamentals including DNS, TCP/IP, and load balancingA passion for automation, continuous improvement, and operational excellenceExperience with Infrastructure as Code such as Terraform, SRE principles such as SLIs and SLOs, or regulated environments is beneficial but not essential Responsibilities: Monitor and maintain highly available production platforms running in AWSRespond to and manage production incidents across a 24/7 serviceInvestigate complex technical issues and restore services quickly and effectivelyDevelop automation to reduce manual operational tasks and improve platform resilienceBuild and improve monitoring, alerting, and observability across cloud environmentsWork alongside Software, Platform, Cloud, and Security Engineers to improve reliability and operational excellenceContribute to post-incident reviews and drive continuous service improvementsSupport containerised workloads using Kubernetes and Docker Technologies: AIAWSBashCloudCloudWatchDatadogDockerGrafanaSupportKubernetesLinuxLoad BalancingPrometheusPythonSecuritySplunkTCP/IPTerraformDevOps More:
We are recruiting Site Reliability Engineers to join a global leader in AI-powered customer experience and cloud technology. Following the award of a major government programme, we are expanding our engineering teams to build and support highly secure, cloud-native platforms that deliver sensitive communication services. This is a fully remote role in the UK with a 24/7 shift pattern on a 28-day rota including days and nights. We offer a competitive salary, bonus, and excellent benefits. You will join an engineering-led organisation where reliability, automation, and continuous improvement sit at the heart of the platform, working collaboratively to build resilient cloud services and help shape the future of highly secure production environments.
last updated 31 week of 2026