About this role
Prometheus · CAD Working Group · Zurich or London or San Francisco
The RoleWe’re hiring a Data Scientist and Engineer to create training data for a frontier CAD and mechanical design model. You’ll work directly with researchers to turn large volumes of CAD and engineering assets into model training datasets and figure out which data actually makes the model better.
This is both a scientific and an engineering role. You’ll investigate the data, design experiments, and build the pipelines that turn your findings into better training datasets.
What You’ll DoTransform raw assets into ML-ready datasets. Extraction, parsing, normalisation, deduplication, and filtering. Some transformations will be mundane, some creative; you’ll decide what’s needed.
Understand and improve the training distribution. Analyse dataset composition, identify gaps and quality issues, and work with researchers to test how data selection, filtering, and mixing affect model performance.
Generate and evaluate synthetic data using LLMs and VLMs. Own generation strategies, prompting, and pipelines at scale. Establish whether the resulting data improves coverage and model performance.
Own the pipelines and their infrastructure. Orchestration, storage, versioning, reproducibility, cost, and performance.
Work with researchers as a peer. When a result looks strange, you’re in the conversation about why—and you’ll often be the one who finds the answer in the data.
What We’re Looking ForMust have
3+ years of hands-on experience in both data science and building data pipelines for machine learning.
Strong statistical reasoning and practical experience designing experiments and evaluating their results.
Strong software engineering skills and comfort owning your own infrastructure.
Nice to have
Exposure to CAD, mechanical design, or geometric data.
Why Join UsCollaborate with world-class researchers on groundbreaking projects.
Competitive salary, benefits, and flexible work arrangements.
A mission-driven culture where your impact is visible.