About this role
Role Overview We are seeking a high-potential, hands-on Lead Data & AI Operations Engineer to own and continuously improve the operational health, governance, controls, reliability, and efficiency of our enterprise Data & AI ecosystem. This is a high-impact technical leadership role with end-to-end accountability for Data & AI Operations across the company. The successful candidate will establish the operating model, engineering controls, automation, observability, and governance required to run Data & AI platforms as reliable, secure, and cost-efficient enterprise services. The ideal candidate combines deep Snowflake and Data Engineering expertise with a strong operations and controls mindset. This engineer will also design, build, and deliver technical solutions and platform capabilities required to achieve operational excellence and efficiency goals.
Key Responsibilities · Supported end-to-end Data & AI Operations and Production Support across enterprise data platforms, data pipelines, analytics, BI, and AI/ML workloads, ensuring availability, reliability, performance, and SLA adherence. · Provided day-to-day Snowflake production support and administration, including workload monitoring, query performance analysis, troubleshooting, access/RBAC management, capacity monitoring, and platform health checks. · Supported and enhanced Data Engineering pipelines and workflows, troubleshooting data ingestion, transformation, orchestration, processing, and downstream data delivery issues across production environments. · Supported AI/ML and GenAI workloads in production, including monitoring application and model-related jobs, data dependencies, API integrations, scheduled processes, failures, and overall operational health. · Contributed to AI Operations (AIOps) capabilities by using AI/GenAI tools for incident analysis, log summarization, anomaly identification, troubleshooting assistance, knowledge retrieval, and faster root-cause analysis. · Developed Python scripts, APIs, workflow automation, RPA, and AI-assisted automation to reduce repetitive operational activities, automate health checks and validations, accelerate issue resolution, and improve support productivity. · Supported the implementation of intelligent monitoring and anomaly detection across data pipelines, Snowflake workloads, and AI services to proactively identify failures, performance degradation, unusual patterns, and operational risks. · Assisted in developing automated remediation and self-healing operational workflows for common production issues, reducing manual intervention and improving the Resolution SLA. · Used GenAI-based operational assistants to support troubleshooting, incident summarization, RCA preparation, log analysis, runbook recommendations, and knowledge management activities. · Monitored production data pipelines, ETL/ELT jobs, orchestration workflows, AI workloads, APIs, and platform services, investigated failures, performed impact analysis, and coordinated timely service restoration. · Performed data quality checks, reconciliation, validation, and root-cause analysis to identify data discrepancies and ensure accurate, complete, and reliable data delivery to downstream applications and AI/analytics workloads. · Supported enterprise data platform controls covering data quality, access, security, privacy, metadata, lineage, change management, and production readiness. · Monitored Snowflake and cloud consumption, performance, and utilization, identified inefficient queries and workloads, and supported optimization initiatives to improve performance and control platform costs. · Built and maintained observability, monitoring, alerting, operational dashboards, automated health checks, and proactive notifications across Data and AI platforms. · Managed Incident, Problem, Change, and Release Management activities, including production troubleshooting, service restoration, RCA documentation, change validation, deployment support, and permanent remediation of recurring issues. · Supported DataOps, MLOps, AIOps, and DevOps practices, including CI/CD pipelines, testing, deployment, release validation, version control, monitoring, documentation, and production support. · Worked closely with Data Engineering, Analytics, BI, AI/ML, Architecture, Security, Infrastructure, and business teams to troubleshoot production issues, manage dependencies, and implement platform improvements. · Participated in on-call and production support activities, ensuring critical Data and AI incidents were addressed within agreed SLAs and appropriately communicated to stakeholders. · Identified recurring operational issues and implemented automation, AI-assisted solutions, process improvements, and permanent fixes to reduce manual effort, prevent repeat incidents, and improve production stability. · Contributed to continuous improvement by promoting operational discipline, automation-first practices, documentation, reusable runbooks, knowledge sharing, and Data/AI production support best practices.
What We Are Looking For · 8–12 years of experience across Data Engineering, Data Platforms, Data Ops, Cloud Engineering, or Production Operations, with demonstrated technical leadership. · Deep hands-on Snowflake expertise, including architecture, administration, SQL, performance tuning, workload management, security/RBAC, monitoring, troubleshooting, and optimization. · Strong experience designing and building engineering solutions, not just administering or supporting Data Platforms. · Demonstrated FinOps and cost optimization experience, with measurable outcomes in Snowflake/cloud consumption reduction, workload optimization, cost attribution, and efficiency improvement. · Strong experience building automation using RPA platforms, Python, APIs, workflow automation, and AI/GenAI tools. · Strong expertise with dBT and enterprise ETL/ELT technologies such as Fivetran, Informatica, and Azure Data Factory. · Experience implementing DataOps, CI/CD, observability, data quality, governance, metadata, lineage, and automated platform controls. · Strong understanding of production operations, incident/problem management, RCA, change management, and platform reliability engineering. · Experience with enterprise BI platforms such as Power BI and Looker. · Ability to operate as both a hands-on engineer and technical leader/manager, taking problems from identification through solution architecture, engineering, implementation, and measurable business outcome. NOTES: · Prefer candidates already residing in Bangalore. · Standard Shift Timing is 12noon to 9pm, however this may vary depending on the business requirements. · 3 Days work from office. · Weekend on call support is required.