About this role
Descripción del puesto / Funciones <ul><li><p>Design and execute end-to-end testing strategies specifically tailored for Machine Learning models, Generative AI systems, RAG architectures, and Autonomous Agents.</p></li><li><p>Validate model accuracy, fairness, bias detection, explainability, robustness, and performance across diverse and edge-case datasets.</p></li><li><p>Execute adversarial testing, prompt-injection, jailbreaking, and red-teaming to evaluate prompt robustness and behavioral variations under stress.</p></li><li><p>Validate agentic workflows, including multi-step reasoning paths, state transitions, tool execution, and fallback behaviors during service failures.</p></li><li><p>Evaluate LLM outputs for correctness, grounding, factuality, consistency, safety, and hallucination reduction.</p></li><li><p>Assess vector store behavior, document chunking logic, retriever configurations, and semantic search accuracy.</p></li><li><p>Conduct API, performance, latency, throughput, and concurrency testing on AI inference endpoints and data pipelines.</p></li><li><p>Ensure compliance with AI ethics, data privacy laws, business rules, and insurance regulatory guidelines, maintaining audit-ready test evidence and behavioral reports.</p></li><li><p>Define AI quality KPIs, establish test governance, and build automated testing frameworks integrated into CI/CD pipelines.</p></li><li><p>Collaborate closely with Data Scientists, ML Engineers, SMEs, and DevOps teams while mentoring junior QA engineers and creating reusable test accelerators.</p></li></ul> Requisitos mínimos <p><br></p><ul><li><p><strong>Experience & Specialization:</strong> Proven senior/lead expertise in software quality engineering with a dedicated focus on AI/ML systems and GenAI applications.</p></li><li><p><strong>Programming & Automation:</strong> Advanced proficiency in Python for test automation, data validation, and custom AI testing scripts.</p></li><li><p><strong>GenAI & RAG Ecosystems:</strong> Hands-on experience with GenAI frameworks, vector databases, chunking strategies, and retrieval evaluation.</p></li><li><p><strong>Model Evaluation & Metrics:</strong> Deep understanding of data validation, model evaluation metrics, fairness/bias testing, and drift detection (data and concept drift).</p></li><li><p><strong>API Testing:</strong> Expertise in testing AI services and model endpoints using tools such as Postman, REST Assured, or Python REST clients.</p></li><li><p><strong>DevOps, Cloud & Infrastructure:</strong></p><ul><li><p>Experience with CI/CD pipelines for continuous testing integration.</p></li><li><p>Exposure to cloud platforms hosting AI deployments.</p></li><li><p>Working knowledge of containerization and orchestration environments (e.g., Docker, Kubernetes).</p></li><li><p>Familiarity with Big Data ecosystems for large-scale AI testing.</p></li></ul></li><li><p><strong>Security & Governance:</strong> Experience in AI ethics, compliance testing, observability tools, and security testing for data pipelines and model-serving endpoints.</p></li></ul> Requisitos valorables <ul><li><p><strong>Advanced Red Teaming:</strong> Hands-on experience building automated adversarial test suites and automated synthetic data generation for rare edge cases.</p></li><li><p><strong>Framework Automation:</strong> Direct implementation of specialized LLM evaluation frameworks (e.g., Ragas, DeepEval, TruLens).</p></li><li><p><strong>Observability Setup:</strong> Advanced configuration of AI monitoring dashboards and automated regression testing workflows for retrained models.</p></li></ul> Idiomas English is a must Ubicación Barcelona