About this role
On the Alexa Daily Essentials Science and Analytics team, we build intelligence systems that power AI-driven experiences for customer-facing features used by hundreds of millions of people daily. Our team applies foundation models in novel ways — designing agentic systems that reason, plan, and act autonomously over complex data, and rigorously evaluating their performance so that every architecture decision is grounded in measured quality. We're not actively training foundation models. We're pushing the frontier of how they're applied — through agentic architectures, reinforcement learning, retrieval-augmented generation, and evaluation science. As an Applied Scientist II, you will design agentic AI systems that autonomously detect, diagnose, and act on signals across Alexa's product surface. You'll build the evaluation science that validates these systems, prototype novel AI-powered features that create new value for customers, and apply reinforcement learning to improve agent behavior through outcome signals. You'll own science end-to-end — from rapid prototyping through production deployment — working directly with engineers and product leaders to ship systems that impact how millions of customers interact with Alexa daily. This is a role for a scientist who brings rigor to evaluation, creativity to invention, and pragmatism to delivery. Key job responsibilities - Design and implement agentic AI systems — multi-step reasoning, tool use, planning, and orchestration — that autonomously surface intelligence and take action over complex, evolving data - Apply reinforcement learning techniques to improve agent routing, decision-making, and task resolution quality from outcome signals - Build evaluation and benchmarking frameworks — automated test suites, LLM-as-judge pipelines, competitive benchmarks, and regression detection across model versions and prompt strategies - Prototype and develop novel AI-powered features — information extraction, proactive recommendations, multi-agent collaboration — from concept through production launch - Build RAG pipelines and knowledge retrieval systems that ground agent reasoning in trusted data - Design and execute experiments end-to-end — hypothesis generation, causal analysis, and results interpretation that directly shape product roadmaps - Communicate findings to technical and non-technical audiences — influencing architecture decisions, model selection, and product strategy through rigorous evaluation A day in the life You might start by analyzing reward signal distributions from your latest RL experiment on agent routing — identifying where the policy improved and where exploration is still needed. Mid-morning, you prototype a new cross-feature experience, testing whether extraction models can reliably surface actionable items from customer content. After lunch, you run your eval suite against a new model version — comparing benchmark scores across reasoning depth, tool-use reliability, and output consistency to inform an architecture decision. Later, you pair with an engineer to add a new capability to an agent's tool registry, then validate it doesn't degrade downstream quality. You end the day designing an A/B test for a proactive feature your team is launching. About the team We're a small, high-ownership team building agentic intelligence systems for Alexa — one of the world's most widely used AI products, reaching hundreds of millions of customers across devices, languages, and contexts. We apply foundation models in novel ways: systems that reason over temporal data, learn from outcomes through reinforcement, autonomously detect regressions, and surface insights that previously required weeks of manual analysis. Our scientists prototype new features, benchmark rigorously, and ship from research through production. If you want to build at the frontier of applied AI — agentic systems, reinforcement learning, evaluation science — with real users and real impact at Alexa scale, this is the team.