Now hiring

ML Ops Engineer @ Anaplan

Cardington Street, LondonOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Salary: £52,000 - 92,000 per year

Requirements: Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure.Proven track record of deploying, scaling, and operationalising machine learning models and LLMs in cloud-native production environments.Demonstrated experience managing compute-intensive GPU infrastructure and high-performance computing (HPC) environments.Advanced proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.Experience with vLLM, Ray, MLflow, LangChain / LangSmith, DeepSpeed, or Hugging Face TGI.Solid background in AWS / GCP / Azure, Kubecost, and GPU cost optimisation techniques.Strong skills in Python, Bash, or Go; deep knowledge of Linux kernel tuning and performance monitoring. Responsibilities: Design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our AI-infused scenario planning platform.Work closely with Data Scientists, ML Engineers, and Cloud Infrastructure teams to streamline model training, deployment, and inference while ensuring optimal GPU utilisation, reliability, and cost-efficiency.Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks such as NVIDIA GPU Operator, Slurm, or Ray.Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible.Optimise GPU compute workloads, high-speed networking, and storage for efficient model training and low-latency inference.Build and maintain robust CI/CD and MLOps pipelines for continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines such as Triton Inference Server, vLLM, and TensorRT-LLM.Enable automated model validation and monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/CPU clusters across AWS, GCP, or Azure.Implement auto-scaling strategies, spot instance policies, and dynamic resource allocation to eliminate infrastructure waste.Establish benchmarking and telemetry to track unit economics and throughput for training and serving AI models.Implement end-to-end observability using tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Technologies: AIAWSAnsibleArgoCDAzureBashCI/CDCloudDevOpsDockerGCPGitHubGrafanaHelmIstioSupportJenkinsKubeflowKubernetesLLMLinuxMachine LearningMLflowMLOpsModel TrainingOpenTelemetryPrometheusPythonTerraformvLLM More:

We are Anaplan, a team of innovators focused on optimizing business decision-making through our AI-infused scenario planning and analysis platform so our customers can outpace their competition and the market. Our customers include more than 2,400 global companies, including Fortune 50 organizations such as Coca-Cola, LinkedIn, Adobe, LVMH, and Bayer. We champion a Winning Culture built on diversity of thought, ambitious goals, leadership at every level, and celebrating wins big and small. We support growth through strategy-led, values-based, disciplined execution, and we welcome what makes each person unique. We are hiring for our Platform Engineering team and offer a collaborative environment where you will be inspired, connected, developed, and rewarded. We also maintain a commitment to DEIB and provide reasonable accommodation for candidates and employees with disabilities.

last updated 36 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores