1,434,312open jobs
83,785companies
217,826added this week
Browse all
Salary
≈ $15k – $35k per year (Estimated)
Location
In office (Gurgaon)
Seniority
Middle · 4+ years exp

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 9, 2026. EXL scores B on the Alion truth index.

Overview
Company
Impact
Profile match

EXL

EXL, legally ExlService Holdings, is a data analytics, AI and digital operations company headquartered in New York that runs outsourced business processes and builds data and AI solutions for insurers, healthcare organisations, banks, media and retail companies. Founded in 1999 and listed on Nasdaq, it has more than 60,000 employees across six continents, with large delivery centers in Noida, Gurgaon, Pune, Bengaluru and Chennai and a newer AI hub in Dublin. It hires data scientists and data engineers, GenAI and MLOps engineers, analytics managers, solution consultants and client partners, plus talent acquisition, finance and medical coding staff.

Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programmed rests on.

  • Ingestion development - build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.
  • Mirroring & CDC - implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.
  • Raw layer management - maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.
  • Standardization & transformation - implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.
  • Data quality - implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.
  • Pipeline operations - schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.
  • Performance tuning - optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.

Documentation - produce and maintain source-to-target mappings, transformation logic documentation and lineage records

Skill Area

Specific Requirements

Core Engineering

Python, PySpark, advanced SQL, Delta Lake, distributed data processing

Microsoft Fabric

Data Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts

Data Integration

Batch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats

Data Quality

Validation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation

Modelling

Bronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes

Ops & Governance

Pipeline monitoring, lineage and metadata capture, access controls, technical documentation

Must-Have Qualifications

  • 4+ years hands-on data engineering with strong PySpark and SQL
  • Production experience building ingestion pipelines from multiple heterogeneous sources
  • Working knowledge of Delta Lake and medallion/lakehouse architecture
  • Experience implementing incremental loads and CDC-style processing
  • Experience implementing data quality checks and troubleshooting pipeline failures

Nice-to-Have

  • Microsoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)
  • Exposure to entity/master data standardization (name and address parsing)
  • Familiarity with libraries such as Great Expectations for data quality
  • Experience optimising for Fabric capacity/CU consumption

Key Deliverables Owned

  • Operational ingestion pipelines for all agreed sources
  • Bronze/raw layer with one Delta table per source and CDC retained
  • Standardization and parsing transformation logic
  • Data quality checks, monitoring and exception handling
  • Source-to-target mapping and run documentation

Dual Role / Complementary Skills

Complementary with the Entity Resolution engineering workstream - both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,434,312 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Gurgaon
≈ $22k – $44k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
SQL
Analytics
Tableau
Apply
≈ $21k – $43k per year (Estimated) • In office • 3+ years exp • Chennai
Python
Databases
Databricks
Trino
AI/ML
Spark
DevOps
Amazon S3
Management
Agile
Scrum
Apply
≈ $21k – $42k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Chennai
Python
Scala
AI/ML
LangChain
NLP
TensorFlow
Keras
PyTorch
Hugging Face
Machine Learning
DevOps
GCP
Azure
AWS
Apply
≈ $21k – $42k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Chennai
SQL
Databases
MySQL
MS SQL
Analytics
ETL/ELT
Talend
Pentaho
Fivetran
Apply
≈ $19k – $38k per year (Estimated) • In office • 8+ years exp • Bengaluru
SQL
Management
Agile
Apply
≈ $14k – $34k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Gurgaon
SQL
Databases
Oracle
Analytics
Power BI
Apply
AWS Data Engineer 1 day ago
≈ $16k – $37k per year (Estimated) • In office • 4+ years exp • Gurgaon
Python
Python
pySpark
Databases
Apache Iceberg
Amazon Redshift
AI/ML
Spark
DevOps
AWS
Amazon S3
Analytics
ETL/ELT
AWS Glue
Apply
≈ $18k – $35k per year (Estimated) • In office • Full-Time • Gurgaon • Chennai
AI/ML
LLM
Knowledge Graph
DevOps
GitHub
Apply
Software Engineer, iOS 10 hours ago
≈ $29k – $71k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru • Gurgaon
Swift
Mobile
SwiftUI
Combine
SPM
GCD
CocoaPods
Clean Architecture
Management
Agile
Apply
≈ $16k – $31k per year (Estimated) • In office • Full-Time • Gurgaon • Chennai
AI/ML
LLM
Knowledge Graph
DevOps
GitHub
Apply
See all jobs
This is one of many
1,434,312 more open roles from verified company boards, updated every day.