Now hiring

Databricks Engineer @ EXL

Chennai, Tamil Nadu, INOnsiteFull-timeJob reference 17933
Apply with ResuMinder

Opens on the employer's site

About this role

We are looking for a skilled and passionate Databricks Engineer to design, build, and optimize enterprise-scale data lakehouse solutions on the Databricks platform. The successful candidate will be responsible for creating Databricks pipeline delivering Financial Crime platforms covering Anti-Money Laundering (AML), Know Your Customer (KYC), Customer Risk Assessment (CRA), Sanctions Screening, Transaction Monitoring, Fraud Detection, and Regulatory Reporting

Databricks Platform Engineering

• Design, build, and maintain Databricks workspaces, clusters, and compute pools across dev/test/prod environments. • Configure and manage Databricks Unity Catalog for data governance, access control, fine-grained permissions, and data lineage. • Optimize cluster configurations — instance types, auto-scaling policies, spot/preemptible nodes — for cost and performance. • Implement workspace-level best practices: folder structures, access controls, secret management (Databricks Secrets / Azure Key Vault / AWS Secrets Manager). • Manage Databricks jobs, workflows, and multi-task job orchestration with dependency management.

Delta Lake & Lakehouse Architecture

• Design and implement Delta Lake tables with appropriate partitioning, Z-ordering, and file compaction (OPTIMIZE / VACUUM). • Build Medallion Architecture (Bronze / Silver / Gold) layers for structured data lake organization. • Implement Delta Live Tables (DLT) pipelines for declarative, reliable ETL/ELT with built-in data quality expectations. • Manage schema evolution, table versioning, time travel, and Change Data Feed (CDF) for incremental processing. • Design data lakehouse patterns integrating Delta Lake with external systems (Kafka, ADLS, S3, GCS).

Data Pipeline Development (PySpark / SQL)

• Develop scalable batch and streaming data pipelines using PySpark, Spark SQL, and Delta Lake. • Build structured streaming pipelines for real-time ingestion from Kafka, Event Hubs, and Kinesis into Delta tables. • Write optimized PySpark transformations leveraging broadcast joins, adaptive query execution (AQE), and dynamic partition pruning. • Create reusable transformation libraries, utility frameworks, and pipeline templates for team productivity. • Implement robust error handling, retry logic, and dead-letter queue patterns in production pipelines.

MLflow & AI/ML Workloads

• Set up and manage MLflow tracking servers, experiment registries, and model lifecycle management on Databricks. • Support data scientists and ML engineers in deploying model training and inference workloads on Databricks clusters and GPU instances. • Build feature engineering pipelines using Databricks Feature Store for reusable, versioned ML features. • Enable GenAI workloads — LLM fine-tuning, RAG pipeline development, and vector search (Databricks Vector Search / Mosaic AI). • Implement MLOps practices: model versioning, A/B testing, model serving via Databricks Model Serving endpoints.

Cloud Integration & DevOps

• Integrate Databricks with cloud-native services: Azure Data Lake Storage (ADLS). • Build and maintain CI/CD pipelines for Databricks notebooks and jobs using Azure DevOps, GitHub Actions, or GitLab CI. • Implement Databricks Asset Bundles (DABs) or Terraform for infrastructure-as-code (IaC) deployment of Databricks resources. • Manage data ingestion using Auto Loader, COPY INTO, and partner integrations (Fivetran, dbt, Airbyte). • Monitor pipeline health, cluster utilization, and costs using Databricks system tables and cloud cost management tools.

Governance, Security & Optimization

• Implement row-level security, column masking, and dynamic data views using Unity Catalog policies. • Ensure data quality enforcement using Delta Live Tables expectations and Great Expectations integrations. • Conduct performance tuning — query plan analysis, caching strategies, Photon engine enablement. • Maintain data cataloging, metadata management, and data lineage tracking within Unity Catalog. • Document architecture decisions, runbooks, and operational guides for Databricks workloads.

Education

• Bachelor's or Master's degree in Computer Science, Information Technology, Data Engineering, or related field.

Experience

• 4-6 years of total experience in data engineering or software engineering. • 2+ years of dedicated hands-on experience with the Databricks platform in production environments. • Strong background in big data engineering, cloud data platforms, and distributed computing.

Ready to apply?

Install the ResuMinder extension and we'll auto-fill the application in seconds — no rewriting.

See how your CV scores