About this role
Job location: Remote India About the role: The Data Engineering team is seeking a Data Engineer with expertise in data infrastructure, pipeline development, and scalable data solutions to play a pivotal role in enabling data-driven decision-making across the organization. The successful candidate will have a deep knowledge of data architecture and engineering best practices, and will work closely with cross-functional teams to ensure data is clean, reliable, and accessible. This role requires strong technical skills, a solid grasp of business needs, and the ability to bridge the gap between raw data and actionable insights through robust engineering solutions. ETL/ELT Pipeline Development: Build, and maintain scalable data pipelines using AWS. Implement both batch and incremental load patterns for BI reporting and application data needs. Real-Time Data Streaming: Develop and manage real-time data ingestion pipelines using Kafka. Ensure low-latency, fault-tolerant data flow for critical business workflows. Workflow Orchestration: Build, schedule, and monitor end-to-end data workflows using Apache Airflow. Manage dependencies, retries, and alerting for production DAGs. Data Warehouse Management: Administer and optimize Amazon Redshift clusters including schema design, query performance tuning, distribution/sort keys, and vacuuming to ensure high availability and cost efficiency. Data Quality & Observability: Implement automated data quality checks at ingestion and transformation stages. Define validation rules, build alerting for anomalies and discrepancies, and establish SLAs to ensure stakeholders can trust the data they use. API Integrations: Integrate third-party and internal REST APIs into data pipelines to pull operational and product data into the warehouse. Cloud Cost Optimization: Monitor and right-size data processing and storage resources across S3, EMR, Redshift, EC2, and Lambda. Proactively identify inefficiencies and propose cost-saving improvements. BI & Analytics Collaboration: Partner with the BI team to align data models, preprocessing logic, and Redshift schema design with reporting and dashboard needs.