Now hiring

Sr. Lead - Service Delivery - Observability & DevOps Platforms @ Ntrs

INOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

About Northern Trust

As a global leader in innovative wealth management, asset servicing, asset management and banking services, Northern Trust (Nasdaq: NTRS) is proud to guide the world’s most successful individuals, families, corporations and institutions.

Since 1889, we have aligned our efforts with our three guiding Principles That Endure: Service, Expertise, and Integrity. Together, they reflect the three cornerstones of business conduct which we strive to instill in our employees, whom we call partners, and to provide to our clients and the communities we serve worldwide.

With more than 135 years of financial experience and over 24,000 partners, we serve the world’s most sophisticated clients using leading technology and exceptional service.

Position : Sr. Lead - Service Delivery - Observability & DevOps Platforms

We are seeking an experienced and highly motivated Service Delivery Sr, Lead to lead the operational governance, service delivery, risk management, customer engagement, and continuous improvement of enterprise Observability and Developer Platform services.

This role serves as the critical bridge between Engineering, Operations Support, Infrastructure Teams, Application Teams, Security, Vendors, and Business Stakeholders, ensuring reliable, secure, scalable, and compliant delivery of strategic platform services.

The Service Delivery Manager will be accountable for the operational success, service quality, risk posture, vendor management, capacity planning, and modernization initiatives supporting enterprise observability solutions, event management platforms, and software delivery toolchains.

The ideal candidate will possess extensive experience operating within large-scale, highly regulated environments such as Financial Services, Manufacturing, Healthcare, or similarly compliance-driven organizations.

Scope of Responsibility

Observability & Monitoring Platforms

• Dynatrace

• Elastic / ELK

• Azure Log Analytics

• SCOM (Microsoft System Center Operations Manager)

• ServiceNow Event Management

• Enterprise Alerting Platforms

Developer Platforms & Productivity Tools

• GitHub Enterprise

• GitHub Actions & GitHub Runners

• Azure DevOps

• Sonatype Nexus Repository

• CI/CD Platforms

• Enterprise Desktop Software & Developer Tooling

Key Responsibilities

Service Delivery & Operational Leadership

• Own end-to-end service delivery for Observability, Monitoring, Event Management, and Developer Tooling platforms.

• Ensure operational excellence and stability of services supporting critical enterprise workloads.

• Maintain high-performing operational support functions for global "Run the Business" activities.

• Establish and monitor service level objectives, KPIs, SLAs, OLAs, and operational scorecards.

• Drive service maturity, operational effectiveness, and customer satisfaction across supported platforms.

• Ensure operational procedures, governance controls, and service management processes are consistently followed.

Engineering & Operations Partnership

• Act as the primary liaison between Engineering Teams, Infrastructure Teams, Operations Support, Security, Vendors, and Application Owners.

• Facilitate seamless transition of projects, upgrades, and new capabilities into production support.

• Drive alignment between strategic engineering initiatives and operational support requirements.

• Ensure operational readiness reviews are conducted prior to production releases.

• Champion supportability, observability, resiliency, and operational excellence during platform modernization initiatives.

Major Incident & Escalation Management

• Own and lead enterprise-wide major incident management processes.

• Serve as the escalation owner during high-priority incidents affecting production services.

• Coordinate engineering, infrastructure, support, vendor, and business stakeholders during service disruptions.

• Ensure effective executive communication throughout incident lifecycles.

• Lead post-incident reviews, Root Cause Analysis (RCA), and corrective action planning.

• Drive permanent resolution of recurring operational issues through Problem Management practices.

Risk, Compliance & Governance

• Ensure compliance with enterprise security, regulatory, and operational standards.

• Identify, assess, track, and mitigate operational risks across services.

• Support internal audits, regulatory examinations, compliance reviews, and risk assessments.

• Maintain service governance frameworks, controls documentation, and operational procedures.

• Partner closely with Risk, Security, Compliance, and Audit organizations to address findings and remediation activities.

• Ensure service delivery aligns with organizational governance requirements and regulatory obligations.

Capacity, Availability & Performance Management

• Own capacity planning processes for observability and platform services.

• Ensure future demand from business growth, projects, and modernization initiatives is incorporated into capacity plans.

• Monitor platform utilization, service performance, and infrastructure health.

• Develop forecasting models and operational dashboards to support strategic planning.

• Drive continuous optimization of availability, performance, scalability, and cost efficiency.

Vendor & Third-Party Management

• Manage strategic relationships with vendors and service providers including Microsoft, Dynatrace, Elastic, GitHub, Sonatype, and managed service partners.

• Conduct service review meetings covering:

• Service Performance

• Risk & Compliance

• Security

• Financial Management

• Continuous Improvement

• SLA Adherence

• Ensure vendors meet contractual obligations and agreed service levels.

• Manage vendor escalations, service improvement plans, and remediation programs.

• Evaluate vendor performance from both operational and financial perspectives.

Service Improvement & Platform Modernization

• Identify opportunities to improve service reliability, efficiency, usability, and customer experience.

• Lead service improvement programs across monitoring, alerting, event management, and CI/CD ecosystems.

• Drive adoption of automation, self-service capabilities, observability best practices, and platform engineering principles.

• Develop and execute Service Improvement Plans (SIPs).

• Ensure actions are tracked through completion with measurable business outcomes.

• Promote automation-first and reliability engineering approaches throughout service operations.

Reporting & Executive Communications

• Provide regular and accurate service performance reporting to leadership.

• Deliver executive dashboards covering:

• Availability

• Reliability

• Capacity

• Risk

• Compliance

• Customer Satisfaction

• Operational Trends

• Present service health reviews to senior leadership and governance forums.

• Communicate service impacts, operational risks, and strategic recommendations to stakeholders.

Required Qualifications

• Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipline.

• 12+ years of experience in IT Infrastructure Operations, Service Delivery, Platform Operations, Engineering Operations, or Production Support.

• Extensive experience in highly regulated enterprise environments.

• Proven experience managing enterprise observability and monitoring ecosystems.

• Strong understanding of:

• Dynatrace

• SCOM

• Elastic / ELK

• Azure Log Analytics

• ServiceNow Event Management

• GitHub Enterprise

• Azure DevOps

• Nexus Repository Manager

• Strong understanding of ITIL disciplines including:

• Incident Management

• Problem Management

• Change Management

• Capacity Management

• Availability Management

• Service Level Management

• Experience managing major incidents and critical production escalations.

• Experience managing vendor contracts and third-party service providers.

• Excellent communication, stakeholder management, negotiation, and leadership skills.

Preferred Qualifications

• ITIL Foundation / ITIL Managing Professional Certification.

• Experience in Financial Services, Manufacturing, Healthcare, or other regulated industries.

• Understanding of Site Reliability Engineering (SRE) principles.

• Experience with ServiceNow ITSM.

• Exposure Cloud Technologies (Azure, AWS).

• Working knowledge of:

• GitHub Actions

• Azure DevOps

• Infrastructure as Code

• Ansible

• Python

• PowerShell

• Automation Frameworks

• Familiarity with Disaster Recovery, Business Continuity, Backup, Storage, Virtualization, Linux, Windows, and Citrix technologies.

Key Competencies

Leadership

• Strategic Thinking

• Executive Communication

• Stakeholder Management

• Vendor Management

• Team Leadership

• Decision Making

Service Management

• ITIL

• Service Governance

• Major Incident Management

• Capacity Management

• Performance Management

• Risk Management

Technical Domain Knowledge

• Observability

• Monitoring

• Event Management

• AIOps

• DevOps Platforms

• CI/CD Tooling

• Platform Operations

Working with Us

As a Northern Trust partner, you will be part of a flexible and collaborative work culture, which has a strong history of financial strength and stability. Movement within the organization is encouraged, senior leaders are accessible, and you can take pride in working for a company committed to an inclusive workplace and assisting the communities we serve.

Philanthropy is deeply rooted in Northern Trust’s history and is an essential element of our culture. Employees around the world give their time and talent to work for the greater good of their communities.

Reasonable Accommodation

Northern Trust is committed to working with and providing adjustments to individuals with health conditions and disabilities. If you need a reasonable accommodation for any part of the employment process, please email our HR Service Center at MyHRHelp@ntrs.com, or alternatively you can discuss your individual requirements with the recruiter you are working with.

About Our Pune Office

The Northern Trust Pune office, established in 2016, is now home to over 3,000 employees. The office handles various functions, including Operations for Asset Servicing and Wealth Management, as well as delivering critical technology solutions that support business operations across the globe.

Our Pune team takes our commitment to service to heart. In 2024, they volunteered more than 10,000+ hours into the communities where they live and work. Learn more.

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores