About this role
DevSecOps-Cloud Infrastructure Engineer Job Description- Cloud Infrastructure Engineer Role Overview We are looking for a Cloud Infrastructure Engineer responsible for designing, building, deploying, and operating enterprise cloud and application infrastructure. The role will own the technical architecture and implementation of servers, virtualization, Kubernetes/OpenShift platforms, networking, storage, and application deployment environments, while working closely with technology vendors and system integrators for implementation and support. Key Responsibilities Design and deploy hybrid cloud and on-prem infrastructure across AWS, VMware, Alibaba Cloud and Tencent Cloud, including GPU compute platforms. Architect and implement Kubernetes/OpenShift platforms, covering cluster design, sizing, HA, networking, storage, security and DR. Define compute, GPU, virtualization, storage, network and capacity architectures for enterprise and cloud-native workloads. Design application platforms using VMs, containers, Kubernetes/OpenShift and cloud-native services; support application onboarding and migration. Implement IaC and automation using Terraform/Ansible and CI/CD pipelines for repeatable infrastructure deployment. Integrate platforms with IAM, DNS, load balancers, PKI, monitoring, logging, backup and security services. Own HLD/LLD, architecture reviews, sizing, implementation plans, acceptance criteria and operational runbooks. Drive performance, resilience, security, capacity management, lifecycle upgrades, patching and technology refresh. Support POCs/MVPs involving cloud, GPU, Kubernetes, automation and emerging infrastructure technologies. Vendor & Partner Management Act as technical owner for cloud, infrastructure and system-integration vendors. Evaluate and recommend solutions across AWS, VMware, Alibaba Cloud, Tencent Cloud, Kubernetes/OpenShift, GPU and virtualization. Review vendor HLD/LLD, configurations, deployment plans and technical deliverables. Opportunities to learning and development and get involvement in the advanced technology projects. The opportunities includes learning and development in advanced technology areas such as-Agentic AI, Knowledge graph, AI security and Quantum Security. Agentic AI & Knowledge Graph: LangGraph, OpenAI/Azure AI Foundry, Neo4j, GraphRAG, vector databases, RAG and AI orchestration frameworks. AI & Quantum Security: Palo Alto AI security, NVIDIA NeMo Guardrails, IBM watsonx.governance, HSM/KMS, PKI, PQC, cryptographic discovery and crypto-agility tools. Technical Skills Strong experience with Linux servers, virtualization, compute, storage, and enterprise infrastructure. Hands-on experience with Kubernetes and/or OpenShift; experience managing vendor-led implementations is highly desirable. Knowledge of private cloud, hybrid cloud, and cloud-native infrastructure. Experience with Docker/container platforms, Helm, Git, and CI/CD. Experience with Terraform, Ansible, or other Infrastructure-as-Code technologies. Strong understanding of networking including DNS, load balancing, routing, firewalls, VPN, ingress/egress, and network segmentation. Knowledge of infrastructure monitoring, logging, observability, backup, and disaster recovery Scripting/programming knowledge in Python, Bash, or similar languages. Qualifications Experience & Qualifications Bachelor's degree in Computer Science, Engineering, Information Technology, or related discipline. 5-8 years of experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or related roles. Proven experience in building enterprise infrastructure and deploying applications. Experience working with technology vendors/system integrators on infrastructure implementation projects. Cloud, Kubernetes, OpenShift, or infrastructure certifications are desirable. Key Competencies Cloud & infrastructure architecture Server and platform engineering Kubernetes/OpenShift platform management Application deployment Infrastructure automation Vendor and system-integrator management DevSecOps and infrastructure security Troubleshooting and technical problem solving Capacity, scalability, availability and resilience Strong technical documentation and stakeholder management