Now hiring

Research Engineer, Pretraining Scaling @ Humanloop

Drew Road, LondonOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

Salary: £57,000 - 73,000 per year

Requirements: We want someone with hands-on experience training large language models, or deep expertise with JAX, TPU, PyTorch, or large-scale distributed systems.We prefer people who genuinely enjoy both research and engineering work, with an ideal split of roughly 50/50.We need someone who is excited about being on-call for production systems, working long days during launches, and solving hard problems under pressure.You should thrive when working on whatever is most impactful, even if priorities change day to day based on the production models needs.We value strong debugging skills across multiple layers of the stack, especially when problems are complex and ambiguous.You should communicate clearly and collaborate effectively, particularly across time zones or during high-stress incidents.We want someone passionate about the work itself and eager to refine their craft as a research engineer.We care about candidates who understand the societal impacts of AI and responsible scaling.A bachelors degree or equivalent combination of education, training, and/or experience is required.Your background should be in a field relevant to the role, demonstrated through coursework, training, or professional experience. Responsibilities: We own critical aspects of our production pretraining pipeline, including model operations, performance optimization, observability, and reliability.We debug and resolve complex issues across the full stack, from hardware errors and networking to training dynamics and evaluation infrastructure.We design and run experiments to improve training efficiency, reduce step time, increase uptime, and enhance model performance.We respond to on-call incidents during model launches, diagnose problems quickly, and coordinate solutions across teams.We build and maintain production logging, monitoring dashboards, and evaluation infrastructure.We add new capabilities to the training codebase, such as long context support or novel architectures.We collaborate closely with teammates across San Francisco and London, as well as with Tokens, Architectures, and Systems teams.We contribute to the teams institutional knowledge by documenting systems, debugging approaches, and lessons learned. Technologies: AIHardwareSupportPyTorchLLMModel TrainingQuant More:

We are Anthropic, a public benefit corporation headquartered in San Francisco, and our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for our users and for society as a whole. Our ML Performance and Scaling team works on training our production pretrained models and operates at the boundary between research and engineering, with deep involvement in performance optimization, hardware debugging, experimental design, and launch coordination. This role requires working in-office 5 days per week in London, and we expect all staff to be in one of our offices at least 25% of the time. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, a lovely collaborative office space, and visa sponsorship support where possible. We are a highly collaborative team that values communication, impact, and high-quality work, and this role offers extraordinary learning opportunities working alongside world-class researchers and engineers on some of the largest training runs in the industry.

last updated 36 week of 2026

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores