Now hiring

Senior AWS Site Reliability Engineers - 2850 @ Xideral

Guadalajara, Jalisco, MéxicoRemoteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Seeking Experienced Senior AWS Site Reliability Engineers for Exciting Projects – Remote in Mexico We are looking for skilled Senior Site Reliability Engineers with a minimum of 5 years of experience to join a dynamic team within a leading organization. This role involves supporting and improving cloud operations for microservice-based platforms, with a focus on production reliability, incident response, cloud infrastructure, automation, observability, Kubernetes operations, and CI/CD workflows across AWS and Azure environments. Key Responsibilities: Own and improve the reliability of cloud-based services and supporting infrastructure. Participate in on-call rotations and support production systems outside normal business hours. Lead incident response activities, including triage, escalation, mitigation, and service restoration. Drive blameless postmortems and ensure corrective actions are tracked to closure. Design, implement, and maintain Infrastructure as Code using Terraform and tools such as Atlantis. Manage and enhance GitOps and deployment workflows using ArgoCD and related CI/CD tools. Support and improve cloud and container platforms across AWS and Azure. Manage Kubernetes-based workloads, containers, virtual servers, and distributed systems. Build automation to reduce manual effort and improve operational efficiency. Configure and improve monitoring, alerting, logging, diagnostics, and observability. Technical Skills Required: With over 5 years of experience as a Senior Site Reliability Engineer, you must be proficient in the following technical skills: Strong hands-on experience with AWS and Azure cloud platforms. Strong experience with Terraform for Infrastructure as Code (IaC). Experience with Atlantis, ArgoCD, or similar infrastructure and deployment automation tools. Strong hands-on experience with Docker and Kubernetes. Experience designing, maintaining, and troubleshooting complex CI/CD pipelines. Strong production support experience, including incident management, Root Cause Analysis (RCA), postmortems, and runbook creation. Strong observability experience, including monitoring, alerting, logging, diagnostics, and performance analysis. Good understanding of cloud networking, security, access controls, and InfoSec practices. Experience with version control, branching, merging, pull requests, and conflict resolution. Understanding of cloud cost optimization and resource utilization. Good-to-Have Skills: Experience with microservice-based platforms. Experience with Datadog, CloudWatch, Grafana, Prometheus, Splunk, AppDynamics, or similar tools. Scripting or programming experience using Python, Bash, Go, or Java. Experience with SLI/SLO/SLA, error budgets, capacity planning, and resilience engineering. Experience with disaster recovery testing and production readiness reviews. Prior experience mentoring junior engineers or leading technical troubleshooting. Qualifications: Bachelor’s degree or higher. Fluent in English (Advanced). Excellent communication, empathy, commitment, leadership, teamwork, and a proactive attitude. Location & Schedule: Remote work from Mexico. Preferred hybrid model in Guadalajara, Jalisco, with expected onsite attendance 2 days per week. Work hours Monday to Friday, 09:00 – 18:00. Advanced English skills are mandatory, and only residents of Mexico. Benefits: Attractive Salary + Premium Benefits Performance bonuses, grocery coupons, and savings are found. Aguinaldo, premium vacations, and vacations paid SGMM Medical insurance, family, and Life insurance. Candidates must include their compensation expectations in their applications and resumes in English. Interested? Apply now through this link:

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores