About this role
AI Evaluation Scientist at Steampunk. Location: McLean, Virginia, United States. Role: implementing evaluation, building pipelines, performing audits Requirements: Ability to design and run AI evaluation frameworks, build automated evaluation pipelines, analyze LLM/RAG behavior, and communicate findings. Requires public trust eligibility, 2+ years evaluating ML/NLP/LLM systems, and proficiency in Python and ML libraries. Category: Data and Analytics Seniority: Entry Level Tools: Python, PyTorch, Hugging Face, scikit-learn, LangChain, Ragas, OWASP LLM Top 10 Certifications: nist ai rmf (aisic), informs cap, aws ml, azure ml, google ml Commitment: Full Time Workplace: Onsite Languages: English