About this role
Salary: £60,000 - 90,000 per year
Requirements: Extensive experience leading Production Support, Site Reliability Engineering (SRE), and Platform Operations functions.Strong expertise in ITIL practices, major incident management, service restoration, problem management, operational resilience, and reliability for critical business services.Strong hands-on knowledge of AWS, microservices, APIs, containerized platforms such as OpenShift and Kubernetes, observability tools, automation, CI/CD, Infrastructure as Code, and modern platform engineering practices.Proven ability to lead critical incidents, conduct root cause analysis, communicate with senior stakeholders, and manage cross-functional teams to deliver high availability, resilience, and positive customer outcomes.Strong understanding of operational risk, controls, audit, governance, regulatory compliance, operational resilience, business continuity, and cloud security in highly regulated environments.Experience managing platform health through SLAs, KPIs, service reviews, capacity planning, proactive monitoring, automation, self-healing capabilities, and continuous improvement initiatives.Demonstrated ability to lead diverse engineering and operations teams, influence enterprise stakeholders, drive operational excellence, foster innovation, and promote modern SRE and cloud-native practices.Relevant critical skills may be assessed, including risk and controls, communication, stakeholder engagement, and role-specific technical skills. Responsibilities: Set the strategic direction for IT Services and implement current methodologies and processes.Manage the IT Services department, including colleague performance, departmental goals, and efficiency and effectiveness.Manage stakeholder relationships and maintain the quality of external third-party services.Develop and implement IT Services policies and procedures, control targets, and standards; ensure adherence to group SLAs and controls for incident, problem, and change activities.Identify and manage IT Services risks, develop mitigation strategies, and maintain alignment with change and compliance functions.Monitor the departments financial performance, control costs, and drive value from commercial agreements.Manage IT Services projects, including research, product launches, and delivery of integrated client solutions.Monitor and maintain critical technology infrastructure, resolve complex technical issues, and minimise operational disruption.Review platform health and performance, oversee incident and problem management, support major incident resolution, and work with engineering teams to maintain highly available, resilient platforms.Coordinate platform releases, drive operational improvements, and ensure services meet reliability and customer experience expectations.Work across Engineering, Infrastructure, Security, and Product teams to automate manual processes, improve monitoring and observability, strengthen resilience, and reduce operational risk.Contribute to service reviews and governance and control activities, and communicate clearly with stakeholders during service events.Advise and influence decision-making, contribute to policy development, and take responsibility for operational effectiveness in collaboration with other functions and business divisions.Lead complex work and assignments, set objectives, coach colleagues, and support performance appraisal and reward decisions where applicable; individual contributors may guide team members and coordinate cross-functional expertise.Consult on complex issues, advise People Leaders on escalated issues, and develop policies and procedures that support risk mitigation, controls, and governance.Analyse information from multiple internal and external sources to solve problems, communicate complex or sensitive information, and influence stakeholders to achieve outcomes. Technologies: AWSCI/CDCloudIncident ManagementSupportITILKubernetesOpenShiftSecuritymicroservicesDevOps More:
We are recruiting a full-time GTSM Senior Service Reliability Engineer within XDP Platforms and Services, based at Knutsford, Radbroke Hall. The role combines technical expertise, operational leadership, and continuous improvement to support secure, reliable, and scalable platform services. We expect colleagues to demonstrate our Barclays Values of Respect, Integrity, Service, Excellence, and Stewardship, and our Barclays Mindset of Empower, Challenge, and Drive. For leadership roles, we expect People Leaders to listen and be authentic, energise and inspire, align across the enterprise, and develop others.
last updated 40 week of 2026