About this role
We are looking for a skilled Data Engineer to design, build, and maintain the infrastructure that supports our data management, processing, and analytics capabilities.
In this role, you will work closely with cross-functional teams to develop scalable data pipelines, modern storage solutions, and effective data governance frameworks. You will play a key role in ensuring that organizational data is accurate, accessible, secure, and optimized for reporting, analytics, and machine learning use cases.
Responsibilities
• Design, develop, and maintain scalable and efficient data pipelines for collecting, processing, transforming, and storing large volumes of data.
• Build, manage, and optimize modern data storage solutions, including data lakes and data warehouses.
• Collaborate with cross-functional teams to define and implement data governance, security, and compliance practices.
• Partner with data analysts and data scientists to optimize data architecture for reporting, analytics, and machine learning workloads.
• Monitor data pipelines and storage infrastructure, identify performance or reliability issues, and implement appropriate solutions.
• Ensure data quality, accuracy, consistency, and integrity through validation, monitoring, and control mechanisms.
• Conduct code reviews and contribute to the continuous improvement of engineering standards and practices.
• Mentor and support junior engineers when required.
• Communicate with business stakeholders to understand data requirements and deliver effective, scalable solutions.
• Stay up to date with emerging technologies, industry trends, and best practices in data engineering.
Requirements
• Bachelor’s degree in Computer Science, Information Systems, Software Engineering, or a related field.
• At least 3 years of hands-on experience in data engineering or a related technical role.
• Strong expertise in SQL, database design, and data modeling.
• Proficiency in at least one programming language, such as Python, Java, or Scala.
• Strong knowledge of ETL and ELT processes, tools, and best practices.
• Experience working with relational and NoSQL databases.
• Experience with Linux-based environments and Bash scripting.
• Hands-on experience with data warehouse technologies such as Snowflake, Amazon Redshift, or Google BigQuery.
• Knowledge of Big Data technologies, including Hadoop and Apache Spark.
• Experience with data streaming technologies such as Apache Kafka and Spark Streaming.
• Experience with workflow orchestration tools such as Apache Airflow, AWS Glue, or similar platforms.
• Practical experience with cloud-based data platforms, including AWS, Microsoft Azure, or Google Cloud Platform.
• Experience implementing CI/CD pipelines for data workflows.
• Familiarity with version control systems, particularly Git.
• Knowledge of data security, privacy, and regulatory compliance standards, including GDPR and HIPAA.
• Exposure to data cataloging, metadata management, and data lineage tools such as Apache Atlas or Amundsen.
• Strong analytical and problem-solving skills.
• Excellent communication and collaboration abilities.
• High attention to detail and the ability to perform effectively in a fast-paced environment.