Now hiring

MTS - Engineering (Data Infrastructure) @ Collinear Ai

San Francisco, California, USOnsiteFull-time
Apply with ResuMinder

Opens on the employer's site

About this role

About the roleAs a Member of Technical Staff - Engineering (Data Infrastructure), you will own the systems that turn large, real-world datasets into data Collinear can build on. Our environments are grounded in terabytes of data, spread across archives, spreadsheets, email, PDFs, and scanned documents. The pace at which we process these determines how fast we deliver to frontier labs.

This is a hands-on role at the intersection of algorithms, distributed systems, and data quality. Many of our hardest problems, such as linking related records across millions of files, don't split up neatly, and you will define how we solve them at scale.

What you'll doBuild pipelines that process multi-terabyte datasets in parallel across archives, spreadsheets, email, PDFs, and scanned documents

Design graph-based systems that link related records, such as the same person or company appearing across millions of files

Build fast string and pattern search over large, heterogeneous datasets

Profile and remove bottlenecks, and decide how to split work that doesn't parallelize neatly

Transform data for downstream use, including consistently replacing sensitive fields across files

Define how we measure data quality, and build review tools so the team can catch and fix errors without reprocessing everything

Assess new datasets and filter out low-quality data before it reaches our environments

About youYou have 5+ years of experience building data-intensive systems in production

You have processed large datasets in parallel with frameworks such as Apache Spark, Ray, or Dask, and know when to design your own

You have a strong command of graph algorithms, and experience using them to transform large amounts of data

You have built efficient string matching and pattern search at scale, such as fuzzy matching or indexing

You make sound tradeoffs between accuracy, speed, and cost, and can explain them clearly

Nice to haveExperience with entity resolution or record linkage

Experience with NLP or LLM-based information extraction

OCR or document processing experience, including poor scans and handwriting

Experience in a systems language such as Rust, C++, or Go

Experience with regulated or sensitive data, such as financial or healthcare records

Skills

Engineering

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores