About this role
Salary: £57,000 - 73,000 per year
Requirements: Significant software engineering experience building and operating production distributed systemsProficiency in at least one systems-appropriate language such as Go, Python, Rust, or C++Deep, hands-on Kubernetes experience beyond basic usage, including scheduler, controllers, apiserver, or operating large multi-tenant clustersDemonstrated ability to debug complex issues across the stack, from API behavior to node- and network-level root causesA track record of designing for reliability, correctness, and clear failure semantics in systems other engineers depend onStrong written and verbal communication skills, with comfort building consensus with internal stakeholdersExperience with Kubernetes internals or contributions such as kube-scheduler, the scheduling framework, apiserver, etcd, client-go, controller-runtime, or similarExperience building or operating cluster schedulers or batch systems such as Kueue, Volcano, Slurm, or in-house equivalentsBackground scaling control planes or coordination systems such as etcd, ZooKeeper, Consul, or large DNS/service-mesh deploymentsFamiliarity with ML infrastructure such as GPUs, TPUs, or Trainium; gang scheduling; topology-aware placement; or collective networking such as NCCLExperience with GCP and/or AWS, including GKE/EKS internals and Infrastructure as CodeLow-level systems experience such as Linux kernel tuning, cgroups, or eBPF12+ years of relevant industry experience, including time leading large, ambiguous infrastructure projectsBachelors degree or an equivalent combination of education, training, and/or experienceA field relevant to the role as demonstrated through coursework, training, or professional experience Responsibilities: Own, operate, and extend the Kubernetes scheduler for our accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane, including apiserver, etcd, and controller-manager, to support clusters far beyond typical limits, and identify the next bottleneck before it finds usDesign, build, and operate core cluster services such as service discovery that every workload in the fleet depends onBuild and maintain custom controllers, operators, and CRDsPartner with research, training, and inference to understand workload shapes and translate requirements into platform capabilitiesCollaborate with cloud providers on required features and escalationsParticipate in on-call, lead incident response, and design processes such as postmortems, runbooks, and SLOs to help the team avoid repeating failures Technologies: AIAPIAWSCloudGCPSupportKubernetesLinuxNetworkPythonRustZooKeeperNodeJS More:
We are Anthropic, a public benefit corporation headquartered in San Francisco, building reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society. Our Kubernetes Platform team runs one of the industrys largest AI compute fleets across multiple cloud providers and datacenters, and we own the control plane that keeps it operating at scale. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office space. We also have a location-based hybrid policy requiring staff to be in one of our offices at least 25% of the time, and we sponsor visas on a case-by-case basis where possible.
last updated 36 week of 2026