Now hiring

Team Lead, DevOps Engineer @ Intermedia Intelligent Communications

United Kingdom, United KingdomRemoteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

*ALL CANDIDATES MUST BE LOCATED IN THE UNITED KINGDOM* About Intermedia: Intermedia has established itself as a leading provider of cloud communications and collaboration tech that allows companies to connect better. We have a strong track record of growth, profitability, and creating an environment where everyone matters. Everyone. While we are fast-paced and admittedly a bit intense, we promise that you won’t be bored. You will find Intermedia is a place where you can indulge your passion for creating and supporting great cloud technology. What’s more, we always look to promote from within and have many employees who have been with us 10, 15, and 20+ years! Are you looking for a company where YOUR VOICE is heard? Where you can MAKE A DIFFERENCE? Do you THRIVE in a FAST-PACED work environment? Do you wake every morning EXCITED to work with GREAT PEOPLE and create SUCCESS TOGETHER? Then Intermedia is the place for you. Culture at Intermedia is built on teamwork and transparency. We hold each other accountable and always have each other’s back! Are you ready to dive into the world of DevOps and make a tangible impact? About the Role: The Unite Services & Unite Integrations teams build and operate the backend platforms that power the Unite app (messenger) and integrations ecosystem. This includes: Integration APIs and microservices Multi-region Kubernetes clusters (cloud & on‑prem) CI/CD and GitOps tooling (GitHub-centric) Internal Developer Platform (IDP) components Critical shared infrastructure (RTE, Redis, Harbor, Consul, Chart Museum, API Gateway, etc.) We are looking for a Team Lead DevOps Engineer to take over the hands‑on technical and leadership scope: owning reliability, security, and operational efficiency for Unite Services & Unite Integrations, while leading a small DevOps team. You will be the go‑to person for Kubernetes, GitHub-based CI/CD, disaster recovery, security hardening, and platform-wide DevOps practices. Team Leadership & Ownership Lead and grow a small DevOps engineering team supporting Unite Services and Unite Integrations. Drive team planning and execution across operational efficiency, security, reliability, and platform maturity epics (e.g., quarterly OKRs/initiatives). Provide technical mentorship on Kubernetes, GitHub, CI/CD, and cloud infrastructure. Collaborate closely with Engineering, SRE, Security, and Ops on roadmap, incident resolution, and cross-team initiatives. Platform Operations (UST / Unite Services / Integrations) Own the operational health of Unite Services and Unite Integrations: Lead production deployments and release processes for backend services (e.g., Unite Notifications, integrations platform, Salesforce-related services). Oversee multi-environment support (DEV/QA/PROD and special environments like RTE). Kubernetes, Cloud & On‑Prem Infrastructure Operate and evolve Kubernetes clusters (cloud and on‑prem), including: Cluster upgrades and migration playbooks Ingress/Nginx and service mesh / networking changes Resource/capacity optimization (CPU throttling, scaling, etc.) Manage supporting infrastructure for Unite platforms: Redis, Consul, Harbor, Chart Museum, API gateways, message queues (e.g., RabbitMQ) Contribute to and adopt Internal Developer Platform (IDP) capabilities (e.g., API Gateway, Apache Flink management, Crossplane-based infrastructure). CI/CD, GitOps & Tooling Lead migration and consolidation of repos and pipelines to GitHub (from TFS/Azure DevOps and others). Design, implement, and maintain GitHub-based CI/CD: GitHub Actions workflows Self-hosted GitHub runners (including K8s-based and Windows runners) Versioning and branching strategies Introduce and maintain GitOps practices for infrastructure and application deployments. Automate repetitive tasks and operational workflows via scripting and infrastructure-as-code. Security, Compliance & Reliability Drive security initiative epics (e.g., Ransomware preparedness, vulnerability remediation like Redis Lua RCE, tagging in Qualys/CrowdStrike). Work with Security to design and implement: Rate limiting, WAF rules, and API protection (e.g., Cloudflare for SCIM and public APIs) Secrets management and secure configuration Define and execute disaster recovery and ransomware recovery strategies for UST / Unite Services / Integrations. Improve reliability and quality of platforms via SLIs/SLOs, operational runbooks, and post-incident improvements. Observability, Incident Management & Runbooks Own observability for Unite Services and Unite Integrations: Metrics, logs, traces (e.g., Prometheus, Grafana, ELK, etc.) Dashboards for capacity (Kubernetes CPU usage, throttling, etc.) and error budgets Lead incident response for platform issues: Triage, mitigation, communication, and follow-up actions Coordination with product/engineering teams during outages Create and maintain detailed runbooks and recovery procedures, especially for: RTE environment, K8s clusters, Harbor, Consul, Chart Museum Disaster recovery and ransomware scenarios Documentation & Process Produce and maintain clear, actionable documentation: Infrastructure design and configuration CI/CD and GitOps processes Recovery and maintenance procedures RFCs for new platform capabilities (e.g., API Gateway, Apache Flink) Contribute to process improvements in CI/CD, deployment strategies, and developer experience. Help define and maintain product maturity models for Unite Integrations and related platforms.

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores