About this role
Responsibilities:
• Administrate the full Kubernetes platform life cycle to ensure the platform remains secure, reliable and highly available. • Work with other engineers to support the infrastructure of OpenShift as well as assist to resolve infrastructure queries/issues of OpenShift tenant applications. • Automation of platform operations and management. • Configure and Manage Monitoring, alerts for the OpenShift environments. • Manage capacity for Kubernetes environments. • Hands-on experience with Docker containerization, Kubernetes orchestration and tools such as Jira and Jenkins. • Hands-on experience in handling Linux OS such as Patching, Filesystem management, User management, SELINUX etc. • Administration and support Windows Servers and Linux server environments. • Perform server hardening, configuration, patching, upgrades, and maintenance. • Troubleshoot IIS, SSL certificate, OpenSSH, tectia and IBM CD file transfer • Monitor server health, CPU, memory, disk, services, and system availability. • Troubleshoot OS, application, network-connectivity, and performance issues. • Handle Linux services, packages, file permissions, SSH, cron jobs, and system logs. • Respond to incidents, alerts, service requests, and production changes. • Perform vulnerability remediation and OS hardening. • Maintain operational documentation, SOPs, and troubleshooting guides.
Requirements:
• 3 years of work experience with a bachelor’s degree in computer science or related Preferred Qualifications. At least 1 year’s hands-on experience with containers in Production Environments - Docker, OpenShift, Kubernetes, Linux and Window Servers preferred • Provide operational support for OpenShift Container Platform (OCP), Red Hat Linux, Windows Server, and infrastructure platforms across production and non-production environments. • Deliver 24x7 operational support for infrastructure and platform services, ensuring timely incident resolution and service restoration. • Perform system administration, monitoring, troubleshooting, performance tuning, and capacity management for Linux, Windows, and containerized environments. • Participate in incident management, problem management, root cause analysis (RCA), and post-incident reviews to improve operational stability. • Maintain operational documentation, standard operating procedures (SOPs), knowledge articles, and support runbooks. • Support security and compliance requirements by executing vulnerability remediation, patch management, access control, and platform hardening activities. • Experience with configuration management tools (Chef, Ansible, terraform etc.). • Hardening, securing the Kubernetes cluster with monitoring and auditing dashboards • Knowledge in infrastructure technologies such as HP and DELL hardware (Blades and Rack servers) • Excellent verbal, written, skills; in particular, demonstrated ability to effectively communicate technical and business issues and solutions to multiple organizational levels internally and externally. • Candidate must have demonstrated and be prepared to exhibit initiative and ownership of consistent delivery success • Be scheduled On-Call to support the infrastructure and systems
Location: DBS Asia Hub Job: Technology Schedule: Regular Employee Status: Full time