About this role
Datavail is a leading provider of data management, application development, analytics, and cloud services, with more than 1,000 professionals helping clients build and manage applications and data via a world-class tech-enabled delivery platform and software solutions across all leading technologies. For more than 17 years, Datavail has worked with thousands of companies spanning different industries and sizes, and is an AWS Advanced Tier Consulting Partner, a Microsoft Solutions Partner for Data & AI and Digital & App Innovation (Azure), an Oracle Partner, and a MySQL Partner. Job Title: Senior Associate Cloud SRE - (Kubernetes, Docker & AWS) Experience: 5 to 10 Years Education: Any graduate Location: Mumbai (Hybrid - Model) Job Summary: We are looking for a highly skilled DevOps / SRE Engineer with strong expertise in Kubernetes, Docker, CI/CD pipelines, and AWS cloud services. The ideal candidate should be capable of troubleshooting complex deployment, infrastructure, and access-related issues across containerized and cloud environments. Key Responsibilities
• Deploy, manage, and troubleshoot applications on Kubernetes clusters
• Diagnose and resolve issues related to nodes, pods, networking, and cluster performance
• Build, manage, and optimize Docker container images
• Manage container image repositories using Harbor (image lifecycle, access control)
• Monitor and troubleshoot CI/CD pipeline failures and deployment issues
• Troubleshoot access and permission issues across DevOps tools and AWS services
• Manage secrets securely using Vault
• Use CLI tools for debugging, automation, and deployments
• Implement and manage infrastructure as code using Terraform
• Support and manage AWS cloud infrastructure and services
Required Skills
• Container & Orchestration
• Strong experience with Kubernetes (K8s) – deployment, scaling, troubleshooting
• Hands-on experience with Docker
• Experience with Harbor (private container registry management)
• CI/CD & DevOps Tools
• Strong knowledge of CI/CD pipelines and troubleshooting
Experience with tools:
• GitHub
• Codefresh
• Artifactory
• Harbor
• TeamCity
• Cloud (AWS)
Hands-on experience with AWS services:
• EC2 – instance management & troubleshooting
• S3 – storage and access control
• IAM – roles, policies, and permission troubleshooting
• ALB (Application Load Balancer) – routing and health checks
• ACM (Certificate Manager) – SSL/TLS certificate management
• Lambda – serverless functions and integrations
• Security Hub – security posture monitoring
• Infrastructure & Security
• Experience with Terraform (IaC)
• Knowledge of HashiCorp Vault for secrets management
• Strong CLI (Linux/Unix) skills
• Ability to troubleshoot access and permission issues across systems
Preferred Skills
• Knowledge of Kubernetes networking (Ingress, Services, DNS)
• Experience with monitoring tools (Prometheus, Grafana, CloudWatch)
• Understanding of container security and vulnerability management
Soft Skills
• Strong analytical and troubleshooting mindset
• Good communication skills for cross-team collaboration
• Ability to work in a fast-paced production environment
• Nice to Have
• Kubernetes certifications (CKA/CKAD)
• AWS certifications
• Exposure to SRE practices (SLI, SLO, Error Budgets)