About this role
About the Role
As an Applied Research Engineer, you’ll build practical AI research assets that support Frontier lab initiatives and customer engagements. This is an implementation-focused role for someone who enjoys turning research concepts into working systems.
You’ll work with a high degree of autonomy, experimenting with new approaches and developing solutions that can be reused across customer opportunities. You’ll partner closely with the GenAI Research team and cross-functional stakeholders to bring technical ideas into practical applications.
Your Impact
Build reinforcement learning and agent environments for real customer and Frontier lab use cases, including task specifications, scoring, and evaluation. Develop benchmarks and evaluation harnesses to measure model and data quality across areas such as accuracy, robustness, safety, latency, and cost. Build LLM pipelines and agentic systems that support research, evaluation, and customer trials. Run fine-tuning, adapter, and other model experiments to evaluate how data and methods influence model behavior. Deploy local or self-hosted models for evaluation, inference, and automation workflows. Document experiments, configurations, data, results, and known limitations so other engineers can reproduce and build on your work. Partner with the GenAI Research team and cross-functional stakeholders to turn technical work into reusable assets for customer engagements.
What You Bring
Bachelor’s, Master’s, or PhD in Computer Science, Engineering, Machine Learning, or a related technical field. 3+ years of professional engineering or relevant industry experience in AI/ML or software engineering. Strong software engineering skills and experience building reliable, maintainable AI systems. Hands-on experience building agentic systems, reinforcement learning environments, LLM pipelines, or similar AI systems. Experience building evaluation harnesses, benchmarks, or model testing pipelines. Ability to work independently on technical problems and move quickly from an idea or research question to a working solution. Strong understanding of experimentation, reproducibility, and technical documentation.
Nice to Haves
Developed synthetic data generation systems or datasets. Published research papers, benchmarks, or other technical research. Worked with SWE-bench or similar software engineering evaluation environments. Built or deployed local inference, open-weight models, or self-hosted model environments.