About this role
At BNY, our culture allows us to run our company better and enables employees’ growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world’s investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide. Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance – and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary. BNY is seeking an accomplished and strategic Senior Vice President – Site Reliability Engineer (SRE) to lead reliability engineering outcomes for critical technology platforms within Global Payment and Trade. This role is designed for a senior engineering leader who combines deep technical expertise with strong execution discipline, influencing the design, resilience, scalability, and operational maturity of large-scale, business-critical systems. In this role, you’ll make an impact in the following ways:
• Lead the strategic direction and execution of site reliability engineering practices across critical platforms, with a focus on resilience, scalability, availability, and operational excellence. • Drive the adoption and maturity of SRE principles, ensuring reliability is engineered into systems from design through production operations. • Define and champion enterprise-grade observability strategies, including monitoring, alerting, logging, tracing, event correlation, and actionable operational intelligence. • Establish, refine, and govern SLIs, SLOs, SLAs, and error budgets to create measurable and business-aligned service reliability objectives. • Lead resilience engineering initiatives, including chaos testing, failure injection, disaster recovery validation, and service hardening, to improve fault tolerance across platforms. • Oversee the identification and elimination of operational toil through automation, self-healing mechanisms, runbook optimization, and platform engineering practices. • Provide leadership during major production incidents, guiding incident response, root cause analysis, post-incident reviews, and long-term corrective actions to prevent recurrence. • Partner with engineering and architecture teams to influence reliability-focused design decisions, ensuring systems are scalable, supportable, and production-ready. • Drive capacity planning, performance engineering, and production readiness assessments for critical applications and services. • Evaluate, recommend, and implement modern tools, frameworks, and engineering practices that improve operational visibility, system health, and reliability outcomes. • Influence and contribute to engineering standards, reliability frameworks, governance practices, and operating models across teams and platforms. • Act as a senior technical leader and trusted advisor, providing thought leadership, mentorship, and technical direction to engineers and engineering leaders. • Build strong partnerships with cross-functional stakeholders to align reliability priorities with business objectives, risk management expectations, and client service outcomes. • Support a culture of continuous improvement, operational accountability, and data-driven decision-making across engineering and support functions. • Drive reliability transformation initiatives that improve MTTR, service availability, change success rate, alert quality, and platform recovery capabilities. To be successful in this role, we’re seeking the following:
• Significant experience in Site Reliability Engineering, Reliability Engineering, DevOps, Platform Engineering, or Production Engineering within complex enterprise environments. • Proven track record of leading large-scale reliability, resilience, and observability initiatives for mission-critical platforms. • Strong expertise in designing and implementing observability solutions using tools such as Splunk, Prometheus, Grafana, Dynatrace, AppDynamics, or similar platforms. • Deep hands-on experience in automation, scripting, and infrastructure as code, using technologies such as Python, Shell, Ansible, Terraform, or equivalent. • Strong experience with chaos engineering, resilience testing, failure scenario design, and service hardening practices. • Excellent troubleshooting and systems-thinking capability across distributed applications, middleware, infrastructure, cloud, and platform services. • Experience with cloud platforms such as AWS, Azure, or GCP, including cloud-native reliability practices. • Strong understanding of Linux/Unix systems, networking, distributed systems architecture, and modern enterprise application landscapes. • Demonstrated ability to lead technical problem-solving across organizational boundaries and influence outcomes at scale. • Strong communication, stakeholder engagement, and executive-level presentation skills. • Experience mentoring engineers and influencing technical direction without necessarily relying on direct line management authority. • Experience supporting or engineering payments platforms, transaction banking systems, or other high-volume, low-latency, highly available environments. • Knowledge of banking, financial services, operational risk, and regulatory expectations related to technology resilience and service continuity. • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline, or equivalent practical experience. Leadership Attributes
• Brings an enterprise mindset, balancing deep technical expertise with strategic business alignment. • Takes full ownership of outcomes and drives execution through complexity and ambiguity. • Influences effectively across engineering, operations, architecture, risk, and senior leadership teams. • Demonstrates sound judgment under pressure, particularly during high-severity incidents and production events. • Champions continuous improvement, engineering discipline, and operational excellence. • Encourages innovation while maintaining strong focus on resilience, control, and sustainable engineering practices.