About this role
Application Monitoring and Operational Support:
• Monitor the health, availability, and performance of applications across production and non-production environments, including development, testing, UAT and staging environments. • Proactively identify application or system issues through monitoring tools, alerts, and operational checks. • Respond to system alerts and operational issues in a timely manner and escalate where required. • Support the ongoing stability and reliability of application environments. Build Deployment and Release Support:
• Deploy application builds across production and non-production environments in accordance with approved release and deployment procedures. • Perform post-deployment verification to ensure successful implementation and application functionality. • Identify and address deployment-related issues and perform rollback procedures where required. • Maintain accurate deployment records and supporting documentation. Log Analysis and Troubleshooting:
• Perform regular reviews of application and system logs to identify errors, exceptions, performance issues, and recurring trends. • Investigate operational issues to determine root causes and appropriate corrective actions. • Escalate technical issues to relevant Development or technical teams where required, providing clear findings and supporting information. • Assist with identifying recurring system issues and opportunities for preventative action. Incident Management:
• Investigate and resolve assigned incidents within agreed service levels and operational procedures. • Maintain accurate and timely updates on incidents and support tickets throughout the resolution process. • Escalate incidents appropriately where further technical intervention is required. • Ensure resolved incidents are closed with appropriate documentation detailing the issue, root cause, and resolution. AI Worker Agents and Automation:
• Design, develop, test, implement, and maintain AI worker agents used to support daily operational processing. • Monitor AI worker agent performance and outputs to ensure accuracy, reliability, and consistent processing. • Investigate and resolve issues relating to automated or AI-enabled processes. • Identify opportunities to automate repetitive operational activities and improve process efficiency. • Support the continuous improvement of existing automation and AI-enabled operational processes. Documentation and Continuous Improvement:
• Develop and maintain operational documentation, including runbooks, deployment procedures, incident records, troubleshooting guides, and AI worker agent documentation. • Ensure operational procedures and supporting documentation remain accurate and up to date. • Identify opportunities to improve operational processes, system reliability, automation, and overall service delivery. • Collaborate with Development and other technical teams to support the resolution of operational issues and implementation of improvements. Standby and After-Hours Support:
• Participate in the Operations standby rotation, currently approximately once per month. • Provide support for critical or urgent operational incidents outside of normal business hours when scheduled for standby. • Respond to and escalate after-hours incidents in accordance with established support and escalation procedures.