About this role
Key Accountabilities/BAU Objectives:
Pipelines from Source to ReportingDesign, build and maintain the pipelines across the medallion layers, from raw ingestion through to the datasets reporting depends on. You own yours end to end, including how they behave when they fail.Pipelines delivered end to end without needing a rebuildRefresh schedules met, with failures caught by monitoring rather than by a userPipelines built to the team standard so someone else can pick them upEngineering Depth and Data Design
Transformations are production code: readable, tested and maintainable. Idempotency and safe replay, correct incremental logic, modelling for how the data is queried, and the judgement to see that generated SQL is the wrong shape even when it returns the right number today.Transformation logic readable, tested and reusable rather than one offPipelines that replay and backfill correctly by designWork another engineer can pick up without a handover conversationIngestion from Product Sources
Bring data in from operational systems through gateways or equivalent connectors, handling incremental loads, schema drift and late arriving data without silent loss.Source data landed completely and repeatedly, reconciled against the sourceIncremental loads correct on replay and on backfillSchema changes detected and handled rather than discovered downstreamData Quality and Schema Consistency
Put quality checks in at each layer and keep schemas consistent across them. Where something is wrong, find where it entered rather than patching the layer it surfaced in.Validation at each layer boundary, with failures visible and ownedDefects traced to the layer they entered and fixed thereA quality check added for every data issue that reached a reportReporting and the Semantic Layer
Build and tune the datasets, models and measures that reporting runs on, and work with analysts and application developers to get them in front of the people who act on them.Reporting backed by fast refresh and metrics that reconcileShared datasets reused rather than duplicated per reportBusiness logic implemented once and shared, not repeated per reportSource Control, Deployment and Support
Work reaches production through source control and a pipeline: small changes, reviewed before merge, nothing edited in place. Monitor what you shipped and support it, including a share of cover for the pipelines your team ownsChanges deployed from source control rather than edited in the portalPipeline failures alerted, triaged and closed to root causeRepeat manual interventions automated away rather than absorbedBuilding with Agents, Checked Against Real Data
Specify the transformation and the result it must produce, direct coding agents to write it, then verify the output against real data before you trust it. Keep the repository context and the checks current so the next person gets the same leverageSpecifications and acceptance criteria written before the buildGenerated logic verified against known data before release, with the check kept as a testRepository context and checks current, and derived from real failuresRequirements
At least 5 years of experience in Data EngineeringUnderstands data architecture. Real delivery against a Medallion or equivalent layered design, with dimensional modelling judgement and an understanding of data warehousing and ETL or ELT designCan build the technology that enables it. Production grade pipelines, advanced SQL including joins, aggregations and incremental loads, and Python or Spark for transformations. Microsoft Fabric is what we run; Databricks, Snowflake or Synapse experience is equally credible.Has run what they built. Pipelines you supported in production, with monitoring, refresh reliability and recovery, and the habit of writing them to be replayed safelyReporting and governance. Power BI dataset modelling, DAX and performance tuning, with metrics that reconcile; data governance, lineage and role based access; and CI/CD for data using Azure DevOps, GitHub Actions or Fabric Git integration.Knowledge and awareness of agentic engineering: what these tools are, where they add value and where their output has to be checked against real data. Hands-on experience is beneficial rather than required.