About this role
Machine Learning - Infrastructure at Causal Labs. Location: San Francisco, California, United States. Role: Design distributed ML training clusters, Develop scalable ML pipelines, Research training approaches Requirements: Design, deploy, and optimize large distributed ML training/inference clusters; develop scalable ML pipelines; research training approaches and optimization strategies. Category: Data and Analytics Seniority: No Prior Experience Required Tools: FSDP, DeepSpeed, Kubernetes, Docker, Google Cloud Platform, Amazon Web Services, Microsoft Azure, Monitoring and observability, Version control, Distributed training frameworks Commitment: Full Time Workplace: Onsite Languages: English