Now hiring

Site Reliability Engineer @ Third Nexus Group Limited

Brayford Square 5, HoveOnsiteFull-timePosted today

Opens on the employer's site

About this role

Salary: £75,000 - 85,000 per year

Requirements: Strong expertise in implementing Site Reliability Engineering (SRE) principles.Advanced knowledge of observability implementation using Dynatrace and Datadog.Proficiency in automation and scripting using Python and Ansible.Strong experience with cloud platforms, including AWS and Azure.Solid understanding of containerization and orchestration tools such as Docker and Kubernetes.Proficiency in cloud-native distributed systems and microservices architecture.Exposure to AI/ML techniques for predictive analytics and automated problem resolution.Familiarity with CI/CD pipelines and automated release and deployment engineering solutions.Experience with chaos engineering tools such as Gremlin or Chaos Monkey, and automation frameworks for resilience tracking.Ability to manage and prioritize multiple projects in a fast-paced environment.Strong interpersonal and communication skills to work effectively across teams.Excellent problem-solving, analytical thinking, and adaptability.Strategic mindset balancing engineering excellence with business priorities.12+ years of experience in IT operations, SRE, or DevOps roles.Proven track record of implementing observability and automation solutions in large-scale environments.Certifications in cloud platforms, observability tools, or other SRE-related areas. Responsibilities: Work closely with the Product Engineering team to modernize IT operations, improve observability, and reduce toil.Architect and deploy observability platforms to monitor system health, performance, and reliability effectively.Propose and drive strategies for AI-driven alerting and proactive anomaly detection to reduce MTTD and MTTR.Develop and enforce SRE best practices, including SLOs, SLIs, and error budgets.Establish and create an AIOps roadmap to improve operational efficiency.Lead efforts to automate repetitive tasks using scripting, orchestration tools, and AI/ML-based solutions.Drive toil automation initiatives for automated incident response and self-healing automation toward autonomous operations.Collaborate with cross-functional teams to ensure systems are scalable, resilient, and maintainable.Drive incident management and root cause analysis processes through automation and continuous improvement.Partner with engineering, architecture, and product teams to enable shift-left engineering practices and improve reliability.Mentor and guide teams on adopting SRE principles and tools.Advocate for a culture of reliability, automation, and continuous improvement across the organization. Technologies: AIAWSAnsibleArchitectAzureCI/CDCloudDatadogDevOpsDockerDynatraceKubernetesPythonmicroservices More:

We are hiring a Site Reliability Engineer for a hybrid role based in Hove, UK, with three days in the office. This is a fixed-term contract offering £85K per annum. You will play a pivotal role in modernizing our IT operations by implementing observability practices, automating toil, and driving reliability at scale. We are looking for a strategic, hands-on leader who can help us build a culture of innovation, resilience, and continuous improvement.

last updated 29 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

Get the extension →
See how your CV scores
Site Reliability Engineer at Third Nexus Group Limited | ResuMinder Jobs