About this role
About Ironsite The construction site is the most dynamic, unstructured physical environment in the world, and until now almost nothing that happens on one has been recorded. Ironsite records it. We design our own smart hard hats with cameras, deploy them with craft workers on the largest infrastructure projects in the country (data centers, LNG facilities, stadiums, hospitals), and turn each shift's footage into actionable insights superintendents read by 5 AM the next morning to improve the speed, efficiency, and predictability of their jobsite. We are pro-worker at our core, empowering the workforce that builds our country. Our footage is now the world's largest dataset of first-person video from active construction sites, and we are using it to build a world model with the spatial intelligence to understand it. Our mission is to create general models that deeply understand dynamic, unstructured environments, predict how work will unfold, and use that understanding to supercharge human productivity and eventually enable robots to fill critical labor gaps. Our culture is one of radical transparency, intellectual honesty, low ego, rapid experimentation, and curiosity. Ironsite is deployed across several of the top ten largest active construction projects in the country. We currently collect ~1,000 hours of new video every day from 8 states and expect to reach ~10,000 hours per day by the end of the year, all while maintaining a worker opt-out rate below two percent, enabled by a workforce-first architecture that anonymizes devices, captures no audio, and never releases raw video. Ironsite is backed by leading investors (8VC, South Park Commons, Saga Ventures) and prominent operators across technology and construction, including Eric Schmidt, Jeff Dean, Jeff Rothschild, Mark Leslie, Scott Wu, Eric Glyman, Karim Atiyeh, Russell Kaplan, and others, alongside over a dozen construction industry operators who have joined us as partners in building this. About the Role As Post-Training Lead , you will report directly to the Chief Science Officer and own our post-training effort end to end: training, benchmarking, and deploying state-of-the-art VLMs that can interpret the complexity of a real-world construction site, built on data no other lab has. This is a hands-on technical lead role: train models, lead the researchers already working on SFT and RL, and help grow the team around this effort. Open problems you could own in your first year: Post-training for expert-level perception. Using SFT and RL (e.g., GRPO) to push VLM labeling of fine-grained construction activity beyond expert human taggers, across every trade. RLVR for agents that investigate the jobsite. Training agents with verifiable rewards to go beyond labeling and run their own research loop over a site's footage: form a hypothesis about what's slowing a crew down, dispatch a swarm to search thousands of hours of video for evidence, and return recommendations that field leaders can act on the next morning. Long-video understanding. Reasoning over multi-hour egocentric footage where the events that matter are sparse. This requires temporal grounding, long-context modeling, and memory beyond what current VLMs offer. From understanding to prediction. Moving from models that describe what happened on a jobsite to models that forecast what happens next: how a crew's current sequence plays out over the coming days, where work will stall, and what changes would prevent it. The first step toward a true world model of construction. Inference at fleet scale. We will soon collect ~10,000 hours of new video per day, expanding quickly from there. Building training and inference systems that scale to all of it at reasonable cost, so frontier models run on every hour we capture. Evals that predict reality. Building benchmarks that track real field performance, which is constantly shifting as we expand to new jobsites, new trades, and even the same jobsite at a different stage of work. What You'll Do Architect & Train Novel VLMs: Design, train, and iterate on general-purpose Vision-Language Models fine-tuned for spatial intelligence in the construction site. Model quality is the single most important output of this role. Own the Post-Training Roadmap: Take the lead on executing our research goals, from establishing baselines with state-of-the-art models to developing post-training recipes (SFT and RL), long-context architectures, and visual reasoning techniques. Build on the Construction Intelligence Benchmark: Expand our existing benchmark suite: video question answering, temporal reasoning, activity recognition, and site-level analytical reasoning. Build Scalable Pipelines: Develop and own the model training and evaluation pipelines, ensuring we can rapidly experiment, measure performance, and deploy models into production. Provide Technical Leadership: Set the technical direction for post-training, review designs and PRs, mentor researchers and interns, and set the standard for experimental rigor. This is technical leadership, not people management. You May Be a Good Fit If You Have owned post-training (SFT and RL) for a large language or vision-language model that shipped to production. Have a background in Computer Science, Machine Learning, AI, Robotics, or a related field. Have demonstrated experience with major deep learning frameworks (e.g., PyTorch, JAX). Are strongly proficient in Python, with a solid foundation in software engineering principles. Have experience working with and creating large-scale vision and/or language datasets. Strong Candidates May Also Have A track record of publications in top-tier AI/ML/CV conferences. Deep expertise in fine-tuning and post-training large language or vision-language models (SFT, GRPO and other RL methods). Experience with technical leadership of other researchers, formally or informally, while continuing to do hands-on research. Hands-on experience with the challenges of video data, such as temporal reasoning, long-context modeling, and efficient processing. Experience optimizing inference, including quantization, distillation, sparsity, and efficient serving. Familiarity with MLOps tools for scalable model training and deployment. A strong interest in vision-language models and applying AI to real-world physical problems, including understanding the day-to-day lives of construction workers. Why You'll Love Working at Ironsite Foundational Impact: Solve fundamental AI problems to transform one of the world's largest and least-digitized industries. Your models ship to real jobsites, not just papers. Ownership & Autonomy: We are a fast-paced startup where you'll have significant ownership over core research directions. We value intellectual curiosity, first-principles thinking, and iterating quickly to turn ambitious ideas into reality. Dream Dataset: Exclusive access to a massive, proprietary, and continuously growing corpus of egocentric jobsite video captured by real construction professionals doing real work on active jobsites. A moat that enables frontier research. World-Class Team: Collaborate with a small, elite team of experts who have a proven track record of building and shipping cutting-edge AI products. Location & Compensation San Francisco Bay Area (on-site) Competitive salary and significant equity package Full benefits including health, dental, vision, and 401k +6% match Access to dedicated GPU compute resources for research and experimentation The base salary range for this role is $250,000 – $400,000 per year.