Now hiring

Site Reliability Engineer | Production @ Snapp

Tehran, Iran, Islamic Republic ofOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

In this role, you will help scale and stabilize our systems as we grow. As part of the SRE team, you'll work on automating operations, managing incidents, and supporting the infrastructure that enables our developers and QA teams to build and release with confidence. You'll be responsible for improving system reliability, monitoring, and observability while ensuring high availability across environments. This role includes participation in a 24/7 shift or on-call rotation. Manage Incidents: Respond to incidents, perform root cause analysis, and help drive resolution and recovery. Monitor & Alert: Improve and tune monitoring systems (Grafana, Prometheus) to ensure issues are detected early. Participate in On-Call: Join a rotating on-call schedule to monitor systems and respond to critical alerts. Collaborate Across Teams: Work closely with developers, QA, and product engineers to support releases and operational improvements. Improve Stability: Proactively identify and fix reliability issues that could affect production uptime. Automate Operations: Build scripts and tools to eliminate manual work and reduce operational overhead. Deploy Services: Assist in deploying and maintaining services across staging and production environments. Support Staging: Troubleshoot and resolve issues in pre-production environments to unblock QA and development teams.

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores