About this role
We are looking for Incident Management Engineers to support a global engagement focused on incident handling, communication, and post-incident reviews. The role requires strong analytical thinking, composure under pressure, and excellent communication skills to manage critical incidents effectively.
• Own the end-to-end lifecycle of all major P1 and P2 incidents within our 16x7 shift window, ensuring response and resolution milestones strictly adhere to corporate SLAs. • Root Cause Analysis: Drive rapid technical triage by analyzing telemetry data, system metrics, monitors and logs to isolate the root cause of complex infrastructure and application failures. • Data-Driven Troubleshooting: Proven ability to quickly interpret telemetry data, consumer lags, and pipeline bottlenecks under high-pressure scenarios to guide engineering teams toward a fix. • Anomaly Mitigation: Actively monitor for performance anomalies, queue lags, and throughput drops to proactively mitigate downstream service degradation. • Tool Proficiency: Experience with observability and monitoring platforms such as Datadog, Grafana
Assertive Communication & Stakeholder Alignment
• Maintain clear, concise, and assertive communication under pressure, cutting through technical noise to extract actionable statuses. • Formulate and broadcast timely business-focused impact statements and progress metrics to executive leadership and client-facing teams. • Isolate technical internal chat channels from high-level notification streams to keep critical updates data-rich and highly orderly.
Post-Incident Evolution & Continuous Improvement
• Perform rigorous Root Cause Analysis (RCA) once an incident is safely stood down. Facilitate and contribute to blameless Post-Incident Reviews (PIR) to track down systematic process or system vulnerabilities. • Actively isolate operational bottlenecks and optimize playbooks to continuously improve Mean Time to Mitigate (MTTM) across critical systems • Maintain structured handoffs between regions (EMEA & APAC)
• Experience in SRE / Incident Management / Production Support • Strong communication & negotiation skills (must-have) • Ability to manage high-pressure situations confidently • Strong problem-solving and analytical mindset • Technical knowledge on task execution. • Good eye for details and understanding of workflows
Technical Skills:
• Ability to proactively identify risks using monitoring tools such as DataDog and Grafana dashboards • Experience in incident response with capability to quickly restore services (restart, patch, or remediate live issues) • Strong focus on minimizing service downtime across environments • Hands-on experience supporting both on-premises (Linux environments) and cloud platforms (primarily Azure, with some exposure to GCP) • Solid understanding of networking concepts and system architecture • Understanding incident impact and skills to analyze and take decisions based on them. • Good understanding of Service now, PagerDuty, JIRA, Databricks, Github actions and basics of Docker and Kubernetes. • Understanding monitoring systems and able to troubleshoot the root cause of issue.