About this role
Design, build, and maintain scalable machine learning (ML) infrastructure to support experimentation, training, deployment, and monitoring of ML models processing large-scale datasets with hundreds of billions of data points. Develop and maintain robust, scalable infrastructure platforms that support the needs of machine learning engineers across multiple business units. Design, build, and maintain data processing and moderation pipelines that handle large data volumes and integrate with trust and safety workflows. Design, develop, and maintain application programming interfaces (APIs), including REST, gRPC, and GraphQL, to support internal ML platform services and system integrations. Develop and implement model evaluation, validation, and quality assurance processes, including A/B testing frameworks and automated evaluation systems, to ensure model accuracy, reliability, and performance. Design, develop, and maintain scalable ML platform systems and data infrastructure using distributed data technologies, including Apache Spark, Kafka, Flink, and Databricks, to support global data processing and analytics needs. Analyze ML infrastructure requirements across business units and design technical solutions within defined scalability, performance, and cost constraints. Support technical design and implementation of ML lifecycle infrastructure, including model training, serving, monitoring, feature stores, and evaluation systems, with an emphasis on platform engineering and self-service capabilities. Participate in hiring activities by conducting technical interviews and providing input on candidate evaluations. Develop and maintain technical documentation, including system designs, operational guides, and internal knowledge bases. Design and optimize recommendation systems and moderation data pipelines, applying best practices for data versioning, feature management, and model evaluation. Implement and optimization of backend and ML services to ensure reproducibility, reliability, and operational stability. Design and optimize large-scale data pipelines and database systems to support efficient data access patterns for ML workflows. Collaborate with cross-functional teams, including software engineers, data engineers, and ML engineers, to support the development and deployment of ML-enabled product features. Design and maintain infrastructure supporting large language model (LLM) workloads. Analyze and resolve complex distributed systems issues affecting performance, scalability, reliability, and availability of high-traffic ML applications. processing jobs, and automation tools supporting the ML lifecycle.