{"id":1247457,"url":"https://alion.io/job/algoworks-lead-data-engineer","title":"Lead Data Engineer","company":{"id":2091826,"name":"Algoworks","domain":"algoworks.com","url":"https://alion.io/company/algoworks-com","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Noida, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":3300000,"max":4000000,"currency":"INR","period":"year","gross":null,"usd_annual":41928},"salary_estimate":null,"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Azure","optional":false},{"name":"Azure Data Factory","optional":false},{"name":"CI/CD","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Git","optional":false},{"name":"pySpark","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Microsoft Fabric","optional":true},{"name":"Python","optional":true}],"status":"live","first_seen_at":"2026-09-24T04:44:07Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-30T00:19:51Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":5},"description":"Role: Lead Data Engineer\nLocation: India, Remote\nExperience: 10+ Years\nAlgoworks\nwww.algoworks.com\nAbout the company\nAlgoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.\nFor over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.\nAt Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.\nThrough collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.\nFollow the video below to know about us! Clipchamp\nRole overview\nWe are seeking a hands-on Senior Data Engineer with strong expertise in Azure Databricks and Azure Data Factory to build, optimize, and maintain scalable enterprise data pipelines.\nThe role will focus on high-performance ETL/ELT development, Delta Lake optimization, data processing across Bronze, Silver, and curated layers, and close collaboration with DWH and reporting teams for downstream consumption.\nThe ideal candidate will act as a senior technical contributor, ensuring reliability, performance, data quality, and maintainability across the data platform.\nKey responsibilities:\n1.Pipeline Development\nBuild and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.\nImplement ingestion and transformation logic across Bronze and Silver data layers.\nDevelop batch and incremental data-processing patterns.\nDesign reliable and reusable pipeline components for enterprise workloads.\nMonitor and troubleshoot pipeline execution and data-processing issues.\n2.Curated Layer & Delta Lake Development\nImplement hydration, merge, and upsert logic using Delta Lake.\nBuild and maintain curated datasets aligned with data quality and business requirements.\nHandle late-arriving data and incremental updates.\nImplement reliable data transformation and reconciliation processes.\nEnsure curated datasets are optimized for downstream consumption.\n3.Performance & Storage Optimization\nOptimize Delta Lake tables for performance and cost efficiency.\nSelect and tune appropriate storage formats such as Parquet and Delta.\nApply partitioning, compaction, and file-sizing strategies.\nTune Spark jobs for large-scale distributed data processing.\nIdentify and resolve performance bottlenecks across data pipelines and storage layers.\n4.Downstream & DWH Collaboration\nWork closely with DWH and reporting teams to support downstream data consumption.\nProvide optimized datasets for reporting and analytical workloads.\nSupport data validation and reconciliation with Gold-layer outputs.\nCollaborate with downstream teams to understand data requirements and optimize delivery.\nEnsure consistency and reliability of data consumed by reporting and analytics platforms.\n5.Engineering Best Practices\nImplement basic CI/CD practices for data pipelines.\nFollow coding standards, documentation, and version-control practices.\nMaintain reusable, scalable, and maintainable pipeline code.\nSupport production troubleshooting and performance tuning.\nParticipate in Agile delivery processes and technical discussions.\n6.Data Quality & Production Support\nImplement data validation and quality checks across ingestion and transformation processes.\nInvestigate data discrepancies and pipeline failures.\nPerform root-cause analysis and implement corrective actions.\nSupport production deployments and resolve data-processing issues.\nMaintain reliability and consistency across enterprise data pipelines.\nRequired technical skills and competencies:\nStrong hands-on experience in data engineering and enterprise data platforms.\nStrong experience building data pipelines on Azure.\nAdvanced proficiency in PySpark.\nHands-on experience with Azure Databricks.\nStrong experience with Azure Data Factory.\nDeep knowledge of Delta Lake tuning and optimization.\nStrong understanding of storage optimization using Parquet and Delta.\nStrong SQL skills for data transformation, validation, and reconciliation.\nExperience working with large datasets and distributed processing.\nExperience implementing batch and incremental processing patterns.\nExperience with hydration, merge, and upsert logic.\nExperience with Git and basic CI/CD pipelines.\nFamiliarity with data quality and validation techniques.\nExperience working in Agile delivery environments.\nMust have skills:\nAzure Databricks.\nAzure Data Factory.\nPySpark.\nDelta Lake.\nSQL.\nData Engineering.\nETL/ELT.\nData Pipeline Development.\nBronze/Silver/Curated Data Layers.\nDelta Lake Performance Optimization.\nSpark Performance Tuning.\nParquet.\nBatch and Incremental Processing.\nGit and Version Control.\nStrong analytical and problem-solving skills.\nGood to have skills:\nMicrosoft Fabric.\nStreaming or near real-time data pipelines.\nData governance tools.\nMetadata management tools.\nAdvanced CI/CD practices.\nExperience with Gold-layer development.\nExperience supporting enterprise reporting and DWH platforms.\nDesired attributes:\nStrong analytical and problem-solving capabilities.\nAbility to independently design and develop complex data pipelines.\nStrong focus on performance, scalability, reliability, and data quality.\nAbility to troubleshoot complex production data issues.\nStrong attention to detail and coding discipline.\nGood communication and collaboration skills.\nAbility to work effectively with DWH, reporting, and cross-functional teams.\nProactive approach to performance optimization and continuous improvement.\nStrong ownership of data engineering deliverables.\nInterview Process\n2-3 rounds of discussion.","description_format":"text","description_chars":6521,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Science & Engineering","Mobile Games","Engineering Services"],"lifecycle":[{"event":"open","at":"2026-09-25T17:58:41Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.858,"p_room":1,"age_days":5,"expected_fill_days":24,"reasons":["conf:11","velocity","win:early"],"computed_at":"2026-09-29T05:45:00Z"},"pay":{"stated_usd_annual":41928,"is_top_pay":false},"html_url":"https://alion.io/job/algoworks-lead-data-engineer","json_url":"https://alion.io/job/algoworks-lead-data-engineer.json","meta":{"generated_at":"2026-09-30T00:33:44Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":385,"day_limit":5000,"remaining_today":4615,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}