Now hiring

Senior Lead Site Reliability / DevOps Engineer @ JP Morgan Chase

Argyle Street 43, GlasgowOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Salary: £62,000 - 102,000 per year

Requirements: We require formal training or certification in software engineering concepts, along with advanced applied experience delivering system design, application development, testing, and operational stability.We require advanced knowledge of reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with in-depth expertise in one or more technical disciplines such as cloud, observability, or distributed systems.We require advanced proficiency in one or more programming languages such as Java, Python, or Go.We require advanced proficiency and experience in observability, including white-box and black-box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, and similar platforms.We require proficiency in continuous integration and continuous delivery tools such as Jenkins, GitLab, and Terraform.We require experience with containers and container orchestration technologies such as ECS, Kubernetes, and Docker.We require hands-on experience designing, deploying, and operating OpenTelemetry collectors in production, including configuring, optimizing, and troubleshooting OTLP endpoints and receivers.We require the ability to solve reliability design and functionality problems independently with little to no oversight.We require practical cloud-native experience.We require the ability to collaborate effectively across different levels and stakeholder groups.Preferred: knowledge of distributed tracing, metrics, and logging best practices.Preferred: certification in AWS, Kubernetes, or relevant technologies.Preferred: proven track record in system health monitoring, capacity management, and blameless postmortems for high-availability services.Preferred: deep understanding of distributed system design principles, networking concepts such as TCP/IP, DNS, and load balancing, and Linux internals.Preferred: contributions to open-source observability or telemetry projects.Preferred: experience with agent control planes and management protocols, with hands-on knowledge of OpAMP highly desirable. Responsibilities: We provide technical guidance and direction on site reliability practices to support our business, technical teams, contractors, and vendors.We develop secure, high-quality production code for reliability tooling and telemetry pipelines, and review and debug code written by others.We drive decisions that influence reliability design, observability architecture, application functionality, and technical operations and processes.We serve as a subject matter expert in one or more areas of site reliability, observability, or telemetry engineering.We lead resiliency design reviews and break complex reliability problems into digestible work for other engineers, acting as a technical lead for large products.We act as the main point of contact during major incidents, identify and solve issues quickly to avoid financial losses, and champion a blameless postmortem culture.We collaborate with team members and stakeholders to define service level indicators, service level objectives, and error budgets.We design, implement, and maintain operational reliability for large-scale OpenTelemetry pipelines in hybrid on-prem and cloud environments, supporting telemetry ingestion, processing, and export to backends such as InfluxDB, Prometheus, Elasticsearch, and OpenSearch.We drive the assessment, refactoring, and incremental migration of custom legacy telemetry collection code to standardized OpenTelemetry instrumentation, reducing technical debt while maintaining system stability.We actively contribute to the engineering community as advocates of firmwide frameworks, tools, and practices, and influence peers and project decision-makers to adopt leading-edge observability and reliability technologies.We contribute to our culture of diversity, opportunity, inclusion, and respect. Technologies: AWSOpenSearchCloudDatadogDockerDynatraceElasticSearchGitLabGrafanaSupportJavaJenkinsKubernetesLinuxLoad BalancingOpenTelemetryPrometheusPythonSecuritySplunkTCP/IPTerraformDevOps More:

We are J.P. Morgan Chase, a global leader in financial services and a leader across banking, markets, securities services, and payments through our Commercial & Investment Bank. We provide strategic advice and products to major corporations, governments, wealthy individuals, and institutional investors in more than 100 countries. Our first-class business in a first-class way approach drives everything we do, and we build trusted, long-term partnerships to help clients achieve their business objectives. We value the diverse talents of our people, are committed to equal opportunity, and foster a culture of diversity, inclusion, and respect. This is a full-time role.

last updated 36 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores