About this role
We are looking for a hands-on Data Engineer with 3+ years of experience building production-grade data pipelines, cloud data platforms, and automated data workflows. You are comfortable working across structured, semi-structured, and unstructured data; you understand the importance of data quality, lineage, security, and cost optimization; and you are excited to build the data foundation required for modern AI, ML, and GenAI use cases.
• Understand business, analytics, and AI use cases and translate them into scalable data engineering solutions. • Design, build, and maintain reliable batch and streaming data pipelines for ingestion, transformation, validation, and publishing. • Develop curated, reusable, and well-documented data products that support BI dashboards, analytics applications, ML models, and GenAI-enabled solutions. • Implement strong data quality checks, observability, lineage, metadata management, and monitoring practices to improve trust in enterprise data assets. • Write clean, modular, and well-tested code using Python, SQL, and modern data engineering frameworks. • Use cloud-native technologies such as BigQuery, Dataflow, Dataproc, Cloud Composer/Airflow, Dataform, DBT, Spark, or equivalent tools to deliver resilient data solutions. • Enable AI/ML and GenAI teams by preparing high-quality feature datasets, vector-ready datasets, document corpora, and governed data access patterns. • Partner with data scientists, ML engineers, product owners, and business stakeholders to support experimentation, model deployment, and production analytics. • Apply DataOps practices including CI/CD, version control, automated testing, reusable templates, release management, and production support standards. • Optimize pipeline performance, storage usage, compute cost, and reliability across cloud-based data platforms. • Support data governance, privacy, access control, and compliance expectations for enterprise and AI-ready data assets. • Stay current with advances in cloud data engineering, AI data infrastructure, orchestration, data quality, and GenAI-enabling technologies.
Minimum Qualifications:
• Bachelor’s or Master’s degree in Computer Science, Data Engineering, Information Systems, Engineering, Statistics, Mathematics, or related technical field. • 3+ years of hands-on experience in data engineering, ETL/ELT development, data warehousing, or cloud-based data platform delivery. • Strong proficiency in SQL and Python for data extraction, transformation, automation, testing, and production support. • Experience designing and operating scalable pipelines on cloud platforms such as Google Cloud Platform, AWS, Azure, or equivalent enterprise data ecosystems. • Experience with modern data platforms and tools such as BigQuery, Spark, Dataflow, Dataproc, Airflow/Cloud Composer, Dataform, DBT, or similar technologies. • Good understanding of data modeling, dimensional modeling, partitioning, clustering, performance tuning, and cost optimization. • Working knowledge of data quality frameworks, monitoring, alerting, metadata, lineage, and production support practices. • Familiarity with Git, CI/CD, agile delivery, code reviews, documentation, and reusable engineering standards. • Strong communication skills with the ability to explain technical solutions clearly to engineering, analytics, and business stakeholders. Preferred Qualifications:
• 5+ years of experience delivering enterprise data engineering solutions in cloud-native environments. • Experience building data products for AI/ML, GenAI, semantic search, retrieval-augmented generation, feature engineering, or model monitoring use cases. • Experience working with unstructured data such as documents, logs, text, images, transcripts, or embeddings, and preparing them for downstream AI consumption. • Hands-on experience with DataOps, MLOps enablement, pipeline observability, automated testing, and production incident resolution. • Experience migrating legacy workflows from Hadoop, Alteryx, or on-premise platforms to modern cloud services. • Experience with APIs, microservices, event-driven architectures, streaming data, or real-time analytics. • Cloud certifications in Google Cloud Platform, AWS, Azure, or relevant data engineering technologies. • Experience mentoring junior engineers, defining engineering standards, or contributing reusable platform accelerators.