Now hiring

Lead Site Reliability Engineer @ Spectrum IT Recruitment

Enterprise Road, SouthamptonOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Salary: £75,000 - 85,000 per year

Requirements: At least six years commercial experience in Site Reliability Engineering or a closely related cloud platform roleDemonstrable experience supporting business-critical cloud platforms and live production servicesStrong hands-on knowledge of Microsoft AzureProduction experience with Kubernetes and containerised workloads, ideally using Azure Kubernetes Service (AKS)Extensive experience in platform engineering, cloud provisioning and observabilityStrong monitoring, alerting and dashboarding experience with technologies such as Azure Monitor, Grafana, Prometheus, OpenTelemetry or ElasticsearchExperience creating custom metrics, queries, dashboards and alerts for microservicesAdvanced scripting or software development skills using PowerShell, Python, C# or a comparable languageStrong Infrastructure as Code experience using Bicep, ARM or TerraformExperience using Git or another version-control platformGood knowledge of Microsoft SQL Server, Elasticsearch and structured data formats, including YAML, JSON and XMLStrong understanding of microservices architecture, cloud platforms and containerisationExperience defining or working with SLOs, SLAs, SLIs and error budgetsExcellent troubleshooting and root-cause analysis skillsExperience designing scalable, secure and maintainable cloud solutionsStrong understanding of cybersecurity principles, governance and complianceExperience working across transformation projects and live-service environmentsApplicants must have lived continuously in the UK for the past five years and be eligible to obtain NPPV3 and UK Security ClearanceDesirable: Azure DevOps pipeline experience covering CI/CD and automated deploymentDesirable: Experience developing reusable infrastructure and monitoring modulesDesirable: Familiarity with AI-enabled engineering and automation toolsDesirable: Knowledge of security and compliance frameworks such as ISO 27001, Cyber Essentials Plus or FedRAMPDesirable: Experience providing technical leadership across multidisciplinary cloud, engineering and support teamsDesirable: A background in regulated, public-sector or security-sensitive environments Responsibilities: Work as part of the Site Reliability Engineering team to protect and improve production environmentsManage and prioritise a technical backlog of reliability, scalability and operational improvementsLead investigations into service outages, performance degradation, platform reliability and cloud expenditureConduct root-cause analysis and ensure corrective actions are implementedIdentify repetitive operational activities and replace them with sustainable automationProvide technical leadership and guidance to Cloud Operations, Support, DevOps and Engineering teamsEstablish and maintain service level objectives, service level agreements, service level indicators and error budgetsDesign and implement monitoring, alerting and dashboards across cloud platforms and microservicesDeploy and configure observability technologies, including Grafana, Prometheus, Azure Monitor and OpenTelemetryDevelop custom application and platform metrics to improve operational visibilityCreate advanced queries, dashboards and alerts for distributed microservicesDevelop reusable Bicep or Terraform modules for monitoring and cloud infrastructureSupport and improve production Kubernetes environments, particularly Azure Kubernetes ServiceReview and optimise platform performance, availability, security and costContribute to cloud architecture, technical scoping and implementation of scalable platform solutionsSupport continuous improvement across deployment, provisioning and operational processesHelp ensure platforms and working practices meet relevant security, governance and compliance requirementsExplore opportunities to use AI-assisted tools to improve automation, troubleshooting and engineering productivity Technologies: AIARMAzureC#CI/CDCloudDevOpsElasticSearchGitGrafanaSupportJSONKubernetesOpenTelemetryPowerShellPrometheusPythonSQLSecurityTerraformXMLmicroservices More:

We provide advanced SaaS solutions to organisations in the public safety and justice sectors, supporting multimedia evidence management and emergency contact centre operations for customers worldwide. As our Cloud Platform Engineering function expands, we are seeking a Senior Site Reliability Engineer for a highly hands-on role focused on keeping business-critical cloud platforms observable, measurable, secure, scalable and reliable. The role is based in Southampton with hybrid working. We welcome candidates from Site Reliability Engineering, DevOps, Cloud Engineering, Platform Engineering or Cloud Development backgrounds. Spectrum IT Recruitment (South) Limited is acting as an Employment Agency in relation to this vacancy.

last updated 40 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores