About this role
Location: India Employment Type: Full-Time Experience: Senior (7+ Years)
Roles & Responsibilities
• Design, build, and maintain scalable batch and real-time data ingestion pipelines.
• Develop and manage Bronze, Silver, and Gold (Medallion) data layers.
• Implement CDC, watermarking, checkpointing, and batch-to-stream data processing.
• Build robust data quality, validation, reconciliation, and monitoring frameworks.
• Develop identity resolution, deduplication, and Golden Record (MDM) solutions.
• Create and maintain source-to-canonical data mappings and crosswalks.
• Ensure schema validation, versioning, and data contract enforcement.
• Collaborate with cross-functional teams to onboard new data sources and optimize data pipelines.
Mandatory Skills
• 7+ years of experience in Data Engineering / Data Pipeline development.
• Strong experience with Apache Kafka (Producers, Consumers, Replay, DLQ, Exactly-once/Idempotent processing).
• Strong SQL and ETL/ELT fundamentals.
• Hands-on experience in Java and/or Python.
• Experience with CDC, Batch & Streaming pipelines, and Medallion/Lakehouse architecture.
• Experience implementing Data Quality, Validation, and Reconciliation frameworks.
• Knowledge of Master Data Management (MDM), Identity Resolution, Deduplication, and Golden Record concepts.
• Experience with source-to-target mapping, canonical data models, YAML/JSON configuration, and Git.
Good to Have
• Experience with probabilistic record matching and record linkage.
• Schema Registry (Avro/Protobuf).
• Experience extracting data from legacy/Mainframe systems.
• Financial reconciliation experience.
• Healthcare or Benefits Administration domain experience.