About this role
Geo Search runs the search that tens of millions of customers use to say where they are going, across many countries where the commercial maps everyone else leans on are often wrong, incomplete, or simply missing. We are building our own search and recommendation stack, which means we own the data, the ranking, and the measurement. The team is cross-functional, spanning backend, machine learning, mobile, and QA, and works with a product manager focused on relevance and with geo analysts. It owns place search and autocomplete, geocoding, ranking and personalization, location detection and pickup selection, and its own search screens on Android and iOS. It is measured on order conversion, search conversion, and mean reciprocal rank We are looking for a Senior Data Scientist to own ranking, relevance and recommendation for geo search: the models that decide which places a customer sees, in what order, and how that changes with context. This is a modeling role with production responsibility, and it is close to a founding one. You will define how relevance is measured here, build the training data from scratch, and own the models that follow. The interesting constraint is that suggestions must return within tens of milliseconds of each keystroke, in cities where map data is thin and people type in mixed scripts, so model choice is an evidence-based trade-off you will own rather than a preference: gradient-boosted ranking on behavioral data today, transformers and large language models where they earn their place in query understanding and cross-script matching, and neural reranking if the failure analysis justifies it. Ground truth comes from what drivers, couriers and customers actually do, and from completed rides, so a large part of the craft is turning messy behavioral logs into labels you can trust. You will measure your impact through offline evaluation and online A/B tests, and changes reach millions of customers in weeks. Own ranking, relevance and recommendation for geo search end to end, from training data through to models serving in production Build the label pipeline that turns search sessions and completed rides into trustworthy training data, including correction for position bias and other presentation effects Train and own the models that order results, and the confidence model that decides per query whether we serve our own answer or fall back to an external provider Build query understanding for our markets, including cross-script matching, using large language models to label offline and distilling into models fast enough for the keystroke path Own evaluation end to end: the offline replay harness that scores recorded sessions against real rides, the error analysis that turns failures into work for the right team, and the online experiments that prove impact Keep models healthy after launch as data and cities shift