About this role
Salary: £57,000 - 73,000 per year
Requirements: We require strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area.We require demonstrated ability to own large experiments end-to-end, from design through interpretation.We require proficiency in Python and experience working with large-scale or distributed ML systems.We require comfort operating at the research/systems boundary, including debugging where the two meet.We require concern for the societal impacts of AI and responsible scaling.We require a bachelors degree or an equivalent combination of education, training, and/or experience.We require a field relevant to the role, as demonstrated through coursework, training, or professional experience.We require years of experience aligned with the internal job level requirements for the position. Responsibilities: We design, run, and interpret large-scale RL experiments, reasoning rigorously about what the data does and does not show.We investigate how RL improves as horizon, compute, and model size grow.We build and maintain benchmarks for long-horizon RL so progress is measurable and reproducible.We translate validated findings into production training recipes, exercising judgment about when a result is robust enough to ship.We debug complex issues at the seam where research meets infrastructure, including failures that only appear at scale.We partner closely with adjacent RL teams across research and engineering and advance our overall RL stack. Technologies: AIPython More:
We are Anthropic, a public benefit corporation headquartered in San Francisco. Our mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for our users and for society. Our team is a quickly growing group of researchers, engineers, policy experts, and business leaders working together on big-science AI research and long-term goals such as steerable, trustworthy AI. We offer competitive compensation and benefits, optional equity donation matching, generous vacation and parental leave, flexible working hours, and a collaborative office environment. We operate with a hybrid policy and expect staff to be in one of our offices at least 25% of the time, and we sponsor visas where possible.
last updated 36 week of 2026