Now hiring

Research Scientist - Vision Foundation Models @ Epsilon Health

San Francisco, California, USOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

About UsWe're tackling one of healthcare's most critical challenges in medical imaging and diagnostics. Our company operates at the intersection of cutting-edge AI and clinical practice, building technology that directly impacts patient outcomes. We've assembled one of the industry's most comprehensive and diverse medical imaging datasets and have a proven product-market fit with a substantial customer pipeline already in place.

Role OverviewWe're seeking a Research Scientist with deep expertise in vision foundation models to join our ML Research team. You'll be at the forefront of developing and deploying state-of-the-art vision models for medical imaging applications. This role focuses on pretraining and scaling vision encoders for radiology diagnosis across X-ray, CT, and MRI, with a growing emphasis on 3D volumetric modeling. You'll work with one of the largest and most diverse medical imaging datasets in the industry, pushing the boundaries of what's possible in AI-assisted diagnosis while maintaining the rigor required for clinical deployment.

Key ResponsibilitiesDesign, train, and scale vision foundation models for radiology applications across X-ray, CT, and MRI modalities, implementing self-supervised, contrastive, masked image modeling, and joint-embedding predictive (JEPA) frameworks.

Extend 2D pretraining recipes to volumetric CT and MR data, addressing long sequence lengths, anisotropic spacing, and multi-sequence studies.

Evaluate model performance rigorously across academic benchmarks, internal offline datasets, and live production data.

Contribute hands-on to all stages of model development including dataset curation, architecture design, distributed training, and production deployment.

Stay current with cutting-edge research in computer vision and medical imaging AI.

Drive research and technical excellence through conference publications and technical blog posts, establishing best practices for training robust medical imaging models at scale.

Qualifications6+ years of academia/industry experience in computer vision/machine learning

Deep expertise in training vision encoder models at scale (e.g. ViT, ConvNeXt). Strong foundation in self-supervised pretraining, including contrastive, masked image modeling, self-distillation, and JEPA-style objectives.

Experience training on volumetric or spatiotemporal data (video, 3D medical imaging)

Track record of implementing complex models from research papers and adapting them to new domains

Proficiency in PyTorch or JAX, with experience training models on multi-GPU/distributed systems

Hands-on experience with medical imaging applications, particularly radiology (X-ray, CT, MRI)

Strong software engineering skills and ability to write production-quality code

Preferred QualificationsPublications at top-tier conferences (CVPR, ICCV/ECCV, NeurIPS, ICLR, MICCAI)

Experience with 3D medical image processing and retrieval tasks

Familiarity with CT and MR acquisition (windowing, multi-sequence protocols, voxel spacing)

Experience with long-context training techniques (sequence parallelism, efficient attention)

Knowledge of vision-language models and multimodal learning

Experience with model interpretability and explainability methods

Understanding of clinical evaluation metrics, clinical workflows, and healthcare data (DICOM, HL7, etc.)

Skills

Research

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores