About this role
As a DevOps & Platform Operations Engineer, you will be responsible for the hands-on technical operations and reliability of the platform, ensuring it is securely deployed, monitored, backed up, and running smoothly in production. You’ll manage deployments, troubleshoot server and integration issues, oversee infrastructure upgrades, and respond to incidents by restoring services and identifying root causes.
The role is focused on keeping the platform secure, available, fast, and reliable while working closely with the product and development functions.
Company Profile:
Our client is an Australian-based claims management company focused on transforming the way insurance claims are managed and settled. Combining extensive industry expertise with technology and dedicated customer service, they deliver efficient, accurate, and transparent claims solutions for insurance companies and their customers.
They are currently seeking for an experienced DevOps & Platform Operations Engineer who is technically capable, detail-oriented, proactive, and energetic, with solid experience in DevOps, cloud infrastructure, platform operations, and maintaining reliable and scalable technology environments.
This is an excellent opportunity to build a long-term career with a growing company, working in a collaborative environment with opportunities to contribute to technology improvements and strengthen platform reliability.
Duties and Responsibilities:
Cloud & Server Management
Manage the SaaS platform’s production and staging infrastructure, including cloud servers and services, databases, domains and DNS, SSL certificates, storage, networking, environment configuration, application services, user access and permissions, and infrastructure scaling. Monitor system health and ensure infrastructure remains appropriately configured as the SaaS platform grows
Deployment & Release Management
Take approved the SaaS platform’s code and reliably deploy it into staging and production environmentsDevelop and maintain automated CI/CD deployment pipelines so releases become increasingly automated and repeatable. Maintain clear separation between Development → Staging → Production. Ensure releases can be rolled back quickly if problems occur
Monitoring & Uptime
Implement monitoring and alerting across the SaaS platform’s environment — application availability, server health, database performance, API performance, error rates, storage, CPU/memory utilization, failed processes, integration failures, security events. Where possible, identify problems before users report them
Incident Response
Act as the first technical point of contact when the SaaS platform experiences an operational issue (website unavailable, application errors, server failures, database problems, failed deployments, API failures, authentication issues, email/notification failures, performance degradation, third- party integration problems). Diagnose the issue, restore service and document the root causeFor significant incidents, produce a short Root Cause Analysis explaining: what happened → why it happened → how it was fixed → how recurrence will be prevented
Backup & Disaster Recovery
Own the SaaS platform’s backup and recovery systems. Ensure databases and critical files are automatically backed up, backups are geographically/reliably stored, recovery procedures are documented, backups are periodically tested, and production environments can be rebuilt if necessary. Maintain a documented Disaster Recovery Procedure
Security
Maintain good cloud and application security practices — access control, multi-factor authentication, secrets management, API credentials, encryption, firewall/security configuration, dependency and vulnerability monitoring, security updates, logging, production access management, backup security. Because the platform handles insurance claim and customer information, security and data protection must be treated as core operational requirementsRequirements
At least 5 years of hands-on experience in DevOps, Platform Operations, Cloud Engineering, or a similar roleStrong experience with Microsoft Azure, including infrastructure, monitoring, networking, security, and production deploymentsExperience with Microsoft Entra ID, access controls, and identity managementExperience managing Linux servers, web applications, and production environmentsExperience with Git/GitHub, CI/CD pipelines, Docker, and application deploymentsKnowledge of DNS, SSL/TLS, SQL databases, backups, REST APIs, and third-party integrationsExperience with cloud monitoring, logging, incident response, troubleshooting, and root-cause analysisUnderstanding of infrastructure and application securityExperience with Python, Bash, PowerShell, or similar scripting languagesStrong problem-solving skills with a methodical and hands-on approachComfortable working independently, taking ownership, and communicating technical issues in plain EnglishComfortable using AI tools as part of everyday workDemonstrates enthusiasm, creativity, and genuine passion for the role and the workShows a high level of engagement, initiative, and ownership in completing tasks and contributing to team objectivesProactive in identifying opportunities, solving problems, and contributing ideas beyond assigned responsibilitiesBrings a positive, collaborative, and results-oriented mindset to the workplace
Advantageous or Nice-to-Have Skills/Experience:
Terraform or other Infrastructure as Code experience