{"id":2164056,"url":"https://alion.io/job/exl-data-engineer-10","title":"Data Engineer","company":{"id":38016,"name":"EXL","domain":"exlservice.com","url":"https://alion.io/company/exl","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Oracle","truth_index":{"grade":"B","score":81,"open_postings":64,"ghost_share":0,"stale_share":0.969,"repost_share":0,"time_to_fill_p50_days":7,"computed_at":"2026-10-09T06:01:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Gurgaon, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":14500,"max_usd":35000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":13},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure Data Factory","optional":false},{"name":"Delta Lake","optional":false},{"name":"Microsoft Fabric","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Great Expectations","optional":true}],"status":"live","first_seen_at":"2026-10-09T07:18:12Z","employer_posted_date":"2026-10-09","last_verified_at":"2026-10-10T01:05:17Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Build and operate the data pipelines that feed the Entity Hub. This role lands all six in-scope sources into Fabric, implements standardization and transformation logic, and maintains the data quality checks and monitoring that the entity resolution engine depends on. Reliable, observable ingestion is the foundation the entire programmed rests on.\n Ingestion development - build and maintain pipelines to land the six in-scope sources (Secretary of State, D&B, ARROW, E1, hCue, DocCentral) into the Fabric Bronze/raw layer.\nMirroring & CDC - implement Fabric Mirroring for supported structured sources and establish change-data-capture patterns; implement watermark/incremental load logic where mirroring is unavailable.\nRaw layer management - maintain one Delta table per source on an append-only basis, retaining evidence records and full source provenance.\nStandardization & transformation - implement name normalization, address parsing and attribute standardization logic in Spark notebooks; support identifier-spine construction.\nData quality - implement data quality checks, validation rules, threshold alerts and exception handling; support reconciliation against source.\nPipeline operations - schedule, monitor and troubleshoot pipeline runs; investigate failures and performance issues; maintain run documentation.\nPerformance tuning - optimise Spark jobs, Delta file sizes, partitioning and pipeline efficiency to manage Fabric capacity consumption.\nDocumentation - produce and maintain source-to-target mappings, transformation logic documentation and lineage records\n\nSkill Area\nSpecific Requirements\n\nCore Engineering\nPython, PySpark, advanced SQL, Delta Lake, distributed data processing\n\nMicrosoft Fabric\nData Factory pipelines and Copy Activity, Lakehouse, OneLake, Spark notebooks, Environments, Mirroring, Shortcuts\n\nData Integration\nBatch and incremental ingestion, CDC patterns, watermarking, reprocessing strategies, schema-on-read for varied formats\n\nData Quality\nValidation rule implementation, completeness/accuracy checks, alerting, exception workflows, reconciliation\n\nModelling\nBronze/Silver/Gold medallion layering, cleansing and conformance, standardization of names, addresses, dates and codes\n\nOps & Governance\nPipeline monitoring, lineage and metadata capture, access controls, technical documentation\n\nMust-Have Qualifications\n4+ years hands-on data engineering with strong PySpark and SQL\nProduction experience building ingestion pipelines from multiple heterogeneous sources\nWorking knowledge of Delta Lake and medallion/lakehouse architecture\nExperience implementing incremental loads and CDC-style processing\nExperience implementing data quality checks and troubleshooting pipeline failures\nNice-to-Have\nMicrosoft Fabric hands-on experience (Mirroring, Copy Jobs, Environments)\nExposure to entity/master data standardization (name and address parsing)\nFamiliarity with libraries such as Great Expectations for data quality\nExperience optimising for Fabric capacity/CU consumption\nKey Deliverables Owned\nOperational ingestion pipelines for all agreed sources\nBronze/raw layer with one Delta table per source and CDC retained\nStandardization and parsing transformation logic\nData quality checks, monitoring and exception handling\nSource-to-target mapping and run documentation\nDual Role / Complementary Skills\nComplementary with the Entity Resolution engineering workstream - both are PySpark-on-Fabric disciplines, so this role can cross-train on Splink tuning and candidate-pair generation to provide cover. Also supports the Sr. Data Engineer (Lead) on identifier-spine construction, and can assist the VectorDB Engineer with document/attribute preparation in Phase 2.","description_format":"text","description_chars":3709,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Business Process Outsourcing (BPO)","AI Consulting & Integration","Analytics & BI Consulting"],"lifecycle":[{"event":"open","at":"2026-10-09T09:04:15Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":7,"reasons":["conf:1","win:early","comp:brand"],"computed_at":"2026-10-10T02:25:38Z"},"pay":null,"html_url":"https://alion.io/job/exl-data-engineer-10","json_url":"https://alion.io/job/exl-data-engineer-10.json","meta":{"generated_at":"2026-10-10T02:25:38Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4458,"day_limit":5000,"remaining_today":542,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":38016},"rest":"https://alion.io/mcp/rest/get_company?id=38016"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fexl-data-engineer-10"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fexl-data-engineer-10"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fexl-data-engineer-10"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/exl-data-engineer-10\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fexl-data-engineer-10"}]}