About this role
Infrastructure SRE - HPC at Sarvam. Location: Bengaluru or Chennai. Role: operating fleet, holding on-call, building tooling Requirements: 5+ years in infrastructure/SRE (including 2+ years operating GPU clusters), on-call ownership and postmortems, Kubernetes fluency, proficiency in Python or Go, and experience with large GPU fleet reliability. Category: Engineering Seniority: Senior Level Tools: Python, Go, Kubernetes, Slurm, NCCL, RDMA, InfiniBand, DGX SuperPOD, MIG, MPS, Lambda, CoreWeave, NeevCloud Commitment: Full Time Workplace: Onsite Languages: English