About this role
Salary: £60,000 - 60,000 per year
Requirements: Ideally have experience in a Production Engineering, Cloud Operations or NOC environmentExposure to Linux systems administrationExposure to AWS cloud infrastructureExposure to Kubernetes and DockerExperience in production support and incident managementScripting experience in Python, Bash or GoExperience with monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatchStrong networking fundamentals including DNS, TCP/IP and load balancingA passion for automation, continuous improvement and operational excellenceExperience with Infrastructure as Code (Terraform), SRE principles (SLIs, SLOs), or regulated environments would be beneficial but is not essential Responsibilities: Monitor and maintain highly available production platforms running in AWSRespond to and manage production incidents across a 24/7 serviceInvestigate complex technical issues and restore services quickly and effectivelyDevelop automation to reduce manual operational tasks and improve platform resilienceBuild and improve monitoring, alerting and observability across cloud environmentsWork alongside Software, Platform, Cloud and Security Engineers to improve reliability and operational excellenceContribute to post-incident reviews and drive continuous service improvementsSupport containerised workloads using Kubernetes and Docker Technologies: AWSBashCloudCloudWatchDatadogDockerGrafanaIncident ManagementSupportKubernetesLinuxLoad BalancingPrometheusPythonSecuritySplunkTCP/IPTerraformNetwork More:
We are hiring a NOC Engineer to help build resilient cloud platforms that support critical national services. This is a fully remote UK role on a 24/7 shift pattern with a 28-day rota including days and nights, offering a competitive salary, bonus and excellent benefits. You will join an engineering-led organisation where reliability, automation and continuous improvement are central to the platform, and where we focus on preventing incidents by improving systems, automating operational processes and shaping highly resilient cloud services.
last updated 31 week of 2026