About this role
<p><span style="font-family:Arial, Helvetica, sans-serif"><span style="font-size:11.0px">At EY, you’ll have the chance to build a career as unique as you are, with the global scale, support, inclusive culture and technology to become the best version of you. And we’re counting on your unique voice and perspective to help EY become even better, too. Join us and build an exceptional experience for yourself, and a better working world for all. </span></span></p> <p> </p> <p> </p> <p> </p> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>EY- Assurance – Senior – Digital</strong></span></p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Role</strong>: <strong>GenAI / Agentic AI Evaluation Engineer (Quality, Safety & Reliability)</strong></span></p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Position Details </strong></span></p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">As part of EY GDS Assurance Digital, you will help design, build, and scale a standardized evaluation capability, focused on evaluating GenAI, RAG-based, and Agentic AI solutions before deployment.</span></p> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">This role sits at the intersection of AI evaluation engineering, Responsible AI, and GenAI security/red teaming. The primary objective is to ensure GenAI/agentic systems are safe, reliable, robust, and fit-for-purpose, by designing evaluation strategies, building repeatable test harnesses, and generating auditable evidence that supports go/no-go decisions.</span></p> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">You will work with global stakeholders (product teams, solution architects, risk & compliance, and assurance leadership) to define evaluation requirements, request test datasets from product teams, execute rigorous evaluations (functional + non-functional), and recommend mitigations and controls to reduce risk.</span></p> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">This is a core full-time role that requires a hands-on AI Development mindset, strong evaluation mindset, and the ability to translate risk concerns into practical testing strategies and measurable acceptance criteria.</span></p> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Responsibilities</strong></span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Define and operationalize evaluation strategies for GenAI systems across use cases like Q&A assistants, summarization, extraction, drafting, agentic systems, and multi-step workflows.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Translate business use-cases into a structured evaluation plan: scope, assumptions, success criteria, datasets, metrics, red-team scenarios, thresholds, and reporting requirements.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Drive standardization: reusable evaluation templates, test case libraries, scoring rubrics, and reporting formats across product teams.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Design structured dataset requirements for product teams and ensure coverage across:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Core user journeys and primary business intents</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Edge cases (rare prompts, ambiguous queries, incomplete context)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Adversarial cases (malicious prompts, jailbreak attempts, prompt injections)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Bias & fairness cases (sensitive demographic proxies, protected attributes, stereotyping patterns)</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Define guidance for dataset sufficiency and statistical coverage (e.g., minimum samples, distribution balance, scenario matrices, stratification by intent/risk).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Build reusable evaluation pipelines for:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Answer quality (correctness, relevance, completeness, clarity)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Grounding & faithfulness (RAG-specific: faithfulness, context precision/recall, hallucination rate, citation quality)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Agentic behavior (tool-call accuracy, tool misuse, goal completion, step correctness, unnecessary actions, loop detection, safety of tool outputs)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Operational quality (latency, cost/token budget, throughput, stability, retries, failure recovery)</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Combine LLM-as-judge and human evaluation in a calibrated way (rubric design, sampling plans, agreement checks).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Implement automated evaluation harnesses in Python (preferred), enabling:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">batch runs on scenario suites</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">configurable metric definitions</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">reproducible runs with run IDs and artifacts</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">storage of traces and outputs for auditability</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Execute structured red teaming aligned to OWASP Top 10 for LLM Applications, covering (examples):</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Prompt injection (direct + indirect) and tool hijacking</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Sensitive data disclosure / PII leakage</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Insecure output handling (downstream injection)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Training data leakage / memorization probes</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Model denial-of-service / denial-of-wallet patterns</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Integrate evals into development lifecycle: pre-release regression gates, CI checks, benchmark comparisons across model versions/prompts/tools/retrievers.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Perform adversarial testing for agentic workflows:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">tool misuse / over-permissioned tool access</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">unauthorized action execution</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">exfiltration via tools/connectors</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">prompt injection via retrieved documents (RAG poisoning)</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Recommend mitigations: input validation, retrieval filtering, tool sandboxing, least-privilege permissions, guardrails, policy prompting, refusal logic, output encoding, monitoring alerts.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Produce high-quality evaluation reports that are auditable and decision-ready, including:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">methodology, datasets, metrics, thresholds</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">quantitative results</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">qualitative results</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">risk assessment summary and recommended control actions</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Present findings to stakeholders in a crisp, risk-informed manner; clearly explain residual risk, limitations, and rationale for go/no-go.</span></li> </ul> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Key Requirements/Skills & Qualification:</strong></span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Excellent academic background, including at a minimum a bachelor’s or a master’s degree in data science, Statistics, Engineering, Operational Research, or other related field with strong focus on modern data architectures, processes, and environments.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">4–7+ years of relevant experience in one or more areas:</span> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">ML/AI/GenAI/Agentic engineering (NLP/LLMs), evaluation engineering, applied research</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">security testing / red teaming</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">building and designing evaluation harness that ensures safety, reliability and robustness.</span></li> </ul> </li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Strong hands-on Python for building evaluation harnesses (data processing, metric computation, orchestration, reporting pipelines).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Practical understanding of GenAI system architectures: RAG, embeddings/vector search, prompt orchestration, tool calling, multi-agent systems, memory, routing.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Experience designing metrics and evaluation methods (rubrics, automated scoring, sampling strategy, regression design).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Familiarity with LLM risks and mitigations, especially for enterprise contexts (data leakage, hallucinations, prompt injection, unsafe content, bias).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Security / Red Teaming Skills (Strong Preference)</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Understanding of OWASP Top 10 for LLM Applications and how to translate it into test cases and controls.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Experience with adversarial testing approaches: jailbreak prompts, injection patterns, tool misuse scenarios, retrieval poisoning patterns.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Familiarity with secure-by-design practices for LLM apps: least privilege, safe tool invocation, output encoding/validation, monitoring.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Evaluation frameworks and tooling: RAGAS, DeepEval, LangSmith, Phoenix/Arize, custom eval harnesses.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Experimentation practices: A/B testing mindset, baseline comparisons, statistical rigor for sample sizes.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Observability/tracing: structured logging, OpenTelemetry, Langfuse-style traces, dashboards.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Basic DevOps practices: Git, CI/CD, containerization (Docker), reproducible environments.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Strong written communication to produce clear evaluation plans and reports for technical + non-technical stakeholders.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Ability to challenge assumptions constructively (“effective challenge”) and influence engineering teams toward remediation.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Comfort operating in ambiguity with fast-evolving GenAI tooling and risk landscape.</span></li> </ul> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Preferred / Nice-to-Have</strong></span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Experience in Assurance/Finance/Regulatory environments (model validation, risk acceptance workflows, audit evidence mindset).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Familiarity with responsible AI frameworks (NIST AI RMF, ISO/IEC 42001, EU AI Act concepts).</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Experience evaluating multilingual systems or domain-heavy enterprise assistants.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Hands-on with Azure ecosystem (Azure OpenAI, AI Search, Function Apps, App Insights, Key Vault).</span></li> </ul> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>Additional skills requirements:</strong></span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Excellent written, oral, presentation and facilitation skills</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Ability to coordinate multiple projects and initiatives simultaneously through effective prioritization, organization, flexibility, and self-discipline.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Must have demonstrated project management experience.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Knowledge of firm’s reporting tools and processes.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Proactive, organized, and self-sufficient with ability to priorities and multitask.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Analyses complex or unusual problems and can deliver insightful and pragmatic solutions.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Ability to quickly and easily create/ gather/ analyze data from a variety of sources.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">A robust and resilient disposition able to encourage discipline in team behaviors</span></li> </ul> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>What we look for</strong></span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">A Team of people with commercial acumen, technical experience, and enthusiasm to learn new things in this fast-moving environment</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">An opportunity to be a part of market-leading, multi-disciplinary team of 7200 + professionals, in the only integrated global assurance business worldwide.</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Opportunities to work with EY GDS Assurance practices globally with leading businesses across a range of industries</span></li> </ul> <p> </p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><strong>What working at EY offers</strong></span></p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">At EY, we’re dedicated to helping our clients, from startups to Fortune 500 companies — and the work we do with them is as varied as they are.</span></p> <p><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">You get to work with inspiring and meaningful projects. Our focus is education and coaching alongside practical experience to ensure your personal development. We value our employees, and you will be able to control your own development with an individual progression plan. You will quickly grow into a responsible role with challenging and stimulating assignments. Moreover, you will be part of an interdisciplinary environment that emphasizes high quality and knowledge exchange. Plus, we offer:</span></p> <ul> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Support, coaching and feedback from some of the most engaging colleagues around</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">Opportunities to develop new skills and progress your career</span></li> <li style="font-family:arial, helvetica, sans-serif;font-size:10.0pt"><span style="font-family:arial, helvetica, sans-serif;font-size:10.0pt">The freedom and flexibility to handle your role in a way that’s right for you</span></li> </ul><p> </p> <p><span style="font-family:Arial, Helvetica, sans-serif"><span style="font-size:11.0px"><b>EY | Building a better working world </b></span></span></p> <p><br> <span style="font-family:Arial, Helvetica, sans-serif"><span style="font-size:11.0px"> <br> EY exists to build a better working world, helping to create long-term value for clients, people and society and build trust in the capital markets. </span></span></p> <p><br> <span style="font-family:Arial, Helvetica, sans-serif"><span style="font-size:11.0px"> <br> Enabled by data and technology, diverse EY teams in over 150 countries provide trust through assurance and help clients grow, transform and operate. </span></span></p> <p><br> <span style="font-family:Arial, Helvetica, sans-serif"><span style="font-size:11.0px"> <br> Working across assurance, consulting, law, strategy, tax and transactions, EY teams ask better questions to find new answers for the complex issues facing our world today. </span></span></p>