About this role
AI Evaluation Scientist at Overview. Location: McLean, Virginia, United States. Role: implementing evaluations, building pipelines, analyzing behavior Requirements: Design and run AI evaluation frameworks, build automated evaluation pipelines, create benchmark datasets, perform LLM/RAG error analysis; requires Python and ML library proficiency and 2+ years evaluating ML or NLP models. Category: Research and Development (R&D) Seniority: Entry Level Tools: Python, PyTorch, Hugging Face, scikit-learn, LangChain, Ragas, OWASP LLM Top 10 Certifications: nist ai rmf (aisic), informs cap, aws ml, azure ml, google ml Commitment: Full Time Workplace: Onsite Languages: English