About this role
Salary: £100,000 - 100,000 per year
Requirements: Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with the ability to implement these practices within an application or platformFluency in at least one programming language, such as Python, Java Spring Boot, or .NETDeep knowledge of software applications and technical processes with emerging depth in one or more technical disciplinesProficiency and hands-on experience in observability practices, including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, or SplunkProficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, or TerraformExperience with container technologies and container orchestration platforms such as ECS, Kubernetes, or DockerExperience troubleshooting common networking technologies and issuesAbility to identify and resolve problems related to complex data structures and algorithmsAbility to collaborate and communicate effectively across different levels and stakeholder groupsExperience mentoring or coaching engineers on site reliability practices and engineering standardsFamiliarity with cloud platforms and infrastructure-as-code practices in large-scale enterprise environmentsExperience contributing to or leading communities of practice, internal knowledge sharing, or engineering guildsExposure to chaos engineering or fault injection methodologies to proactively test system resilienceAbility to evaluate and introduce emerging technologies that improve platform reliability and reduce operational toil Responsibilities: Demonstrate and champion site reliability culture and practices, and exert technical influence throughout our teamLead initiatives to improve the reliability and stability of our teams applications and platforms using data-driven analytics to improve service levelsCollaborate with team members to identify comprehensive service level indicators and establish reasonable service level objectives and error budgets with customersDemonstrate a high level of technical expertise within one or more technical domains and proactively identify and solve technology-related bottlenecks in our areas of expertiseAct as the main point of contact during major incidents for our application and identify and solve issues quickly to avoid financial lossesDocument and share knowledge within our organization via internal forums and communities of practiceTake the lead on resiliency design reviewsBreak up complex problems into digestible work for other engineersAct as a technical lead for medium to large-sized productsProvide advice and mentoring to other engineers Technologies: CloudDatadogDockerDynatraceGitLabGrafanaJavaJenkinsKubernetesMarketingPrometheusPythonSecuritySplunkSpringSpring BootTerraformASP.NETDevOps More:
hackajob is partnering directly with JPMorganChase to hire for this role. We are looking for a Lead Site Reliability Engineer within our Chief Technology Office, where you will help define the future of a globally recognized firm and have a direct and significant impact in a realm tailored for top achievers in site reliability. J.P. Morgan is a global leader in financial services, providing strategic advice and products to the worlds most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do, and we strive to build trusted, long-term partnerships to help our clients achieve their business objectives. We value the diverse talents our people bring to our global workforce, and we are committed to diversity and inclusion, equal opportunity, and reasonable accommodations. Our Corporate Functions team covers a diverse range of areas from finance and risk to human resources and marketing, and plays an essential role in setting our businesses, clients, customers and employees up for success.
last updated 36 week of 2026