About this role
Join our dynamic team to innovate and refine technology operations, impacting the core of our business services.
As a Technology Support Lead in the Technology Support team, you will play a leadership role in ensuring the operational stability, availability, and performance of our production services.
Job Responsibilities
• Own end-to-end production support for mission-critical applications and platform components supporting treasury/trading/risk/data/reporting workflows • Lead major incident (P1/P2) response: triage, decisioning, restoration, communications, and post-incident review • Drive problem management and stability engineering by setting standards for RCA quality and timeliness • Define prevention roadmaps, track actions to closure, and reduce repeat incidents while improving MTTR/MTBF • Apply an SRE approach: identify operational toil, prioritize automation, and deliver measurable reduction in manual effort and failure rates • Build and enforce governance across production operations (runbook standards, operational readiness reviews, change governance, control/evidence routines) • Provide KPI/MIS reporting on incidents, availability, batch health, and risk themes • Strengthen observability and monitoring across applications and data pipelines (define SLOs/SLIs, alerting standards, improve signal-to-noise, dashboards, proactive detection) • Lead batch and data operations governance across AutoSys, Airflow, and data platforms including Databricks/DataLake/ETL • Partner with engineering/platform teams to improve resiliency (capacity, performance, HA/DR, error budgets where applicable) • Coach and lead a support team (where applicable) and act as a senior stakeholder interface for business, technology, and control partners with clear executive-level communication Required Qualifications, Capabilities, and Skills
• 8+ years of experience in production/application support and/or SRE/operations for mission-critical platforms in banking/financial services (mandatory), including leadership accountability • Demonstrated SRE mindset and execution with proven toil identification and elimination • Proven delivery of automation and operational simplification • Strong governance and controls ownership (audit-ready) • Strong hands-on technical depth in Linux and SQL • Automation skills in Python and shell scripting • Scheduling/orchestration experience with AutoSys and Airflow • Cloud/platform experience with AWS, Kubernetes, and cloud technologies • Data platform experience with Databricks, DataLake, and ETL • Observability experience with Grafana, Dynatrace, Splunk, and OpenTelemetry • Strong incident/problem/change management expertise, including operating under ambiguity and time pressure, plus experience with SLOs/reliability dashboards/alert tuning, DR testing, resiliency reviews, operational risk assessments, and senior stakeholder communication/reporting Preferred Qualifications, Capabilities, and Skills
• Experience with containerized microservices, service meshes, and event-driven architectures.