About this role
We are currently actively building a Data Warehouse a key part of the product. We work with cutting edge technologies (GCP, AWS, Airflow, Kafka, K8s) and make infrastructure and architectural decisions based on data. We are building a large scale data infrastructure for analytics, machine learning, and realtime recommendations. Our tech stack Languages: Python, SQL Frameworks: Spark, Apache Beam Storage and analytics: BigQuery, GCS, S3, Trio, other GCP and AWS stack components Integration: Apache Kafka, Google Pub/Sub, Debezium ETL: Airflow 2 Infrastructure: Kubernetes, Terraform Development: GitHub, GitHub Actions, Jira Gather and clarify requirements from diverse stakeholders across the company. Design and evolve DWH Architecture (ODS and Data Mart layers) with a focus on scalability, performance, and data security standards. Build robust and efficient incremental pipelines; develop and optimize data marts in BigQuery (Dataform/SQL/DBT) and Airflow. Participate in testing, data validation, and release processes Design and implement data quality checks; investigate data quality issues and consistency discrepancies across various pipelines. Perform deep-dive analysis of source systems to build efficient data flows from source to consumption. Maintain architectural and technical documentation to ensure data transparency and compliance.