Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 24, 2026.
Role: Lead Data Engineer
Location: India, Remote
Experience: 10+ Years
Algoworks
About the company
Algoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.
For over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.
At Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.
Through collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.
Follow the video below to know about us! Clipchamp
Role overview
We are seeking a hands-on Senior Data Engineer with strong expertise in Azure Databricks and Azure Data Factory to build, optimize, and maintain scalable enterprise data pipelines.
The role will focus on high-performance ETL/ELT development, Delta Lake optimization, data processing across Bronze, Silver, and curated layers, and close collaboration with DWH and reporting teams for downstream consumption.
The ideal candidate will act as a senior technical contributor, ensuring reliability, performance, data quality, and maintainability across the data platform.
Key responsibilities:
1.Pipeline Development
- Build and maintain scalable data pipelines using Azure Databricks and Azure Data Factory.
- Implement ingestion and transformation logic across Bronze and Silver data layers.
- Develop batch and incremental data-processing patterns.
- Design reliable and reusable pipeline components for enterprise workloads.
- Monitor and troubleshoot pipeline execution and data-processing issues.
2.Curated Layer & Delta Lake Development
- Implement hydration, merge, and upsert logic using Delta Lake.
- Build and maintain curated datasets aligned with data quality and business requirements.
- Handle late-arriving data and incremental updates.
- Implement reliable data transformation and reconciliation processes.
- Ensure curated datasets are optimized for downstream consumption.
3.Performance & Storage Optimization
- Optimize Delta Lake tables for performance and cost efficiency.
- Select and tune appropriate storage formats such as Parquet and Delta.
- Apply partitioning, compaction, and file-sizing strategies.
- Tune Spark jobs for large-scale distributed data processing.
- Identify and resolve performance bottlenecks across data pipelines and storage layers.
4.Downstream & DWH Collaboration
- Work closely with DWH and reporting teams to support downstream data consumption.
- Provide optimized datasets for reporting and analytical workloads.
- Support data validation and reconciliation with Gold-layer outputs.
- Collaborate with downstream teams to understand data requirements and optimize delivery.
- Ensure consistency and reliability of data consumed by reporting and analytics platforms.
5.Engineering Best Practices
- Implement basic CI/CD practices for data pipelines.
- Follow coding standards, documentation, and version-control practices.
- Maintain reusable, scalable, and maintainable pipeline code.
- Support production troubleshooting and performance tuning.
- Participate in Agile delivery processes and technical discussions.
6.Data Quality & Production Support
- Implement data validation and quality checks across ingestion and transformation processes.
- Investigate data discrepancies and pipeline failures.
- Perform root-cause analysis and implement corrective actions.
- Support production deployments and resolve data-processing issues.
- Maintain reliability and consistency across enterprise data pipelines.
Required technical skills and competencies:
- Strong hands-on experience in data engineering and enterprise data platforms.
- Strong experience building data pipelines on Azure.
- Advanced proficiency in PySpark.
- Hands-on experience with Azure Databricks.
- Strong experience with Azure Data Factory.
- Deep knowledge of Delta Lake tuning and optimization.
- Strong understanding of storage optimization using Parquet and Delta.
- Strong SQL skills for data transformation, validation, and reconciliation.
- Experience working with large datasets and distributed processing.
- Experience implementing batch and incremental processing patterns.
- Experience with hydration, merge, and upsert logic.
- Experience with Git and basic CI/CD pipelines.
- Familiarity with data quality and validation techniques.
- Experience working in Agile delivery environments.
Must have skills:
- Azure Databricks.
- Azure Data Factory.
- PySpark.
- Delta Lake.
- SQL.
- Data Engineering.
- ETL/ELT.
- Data Pipeline Development.
- Bronze/Silver/Curated Data Layers.
- Delta Lake Performance Optimization.
- Spark Performance Tuning.
- Parquet.
- Batch and Incremental Processing.
- Git and Version Control.
- Strong analytical and problem-solving skills.
Good to have skills:
- Microsoft Fabric.
- Streaming or near real-time data pipelines.
- Data governance tools.
- Metadata management tools.
- Advanced CI/CD practices.
- Experience with Gold-layer development.
- Experience supporting enterprise reporting and DWH platforms.
Desired attributes:
- Strong analytical and problem-solving capabilities.
- Ability to independently design and develop complex data pipelines.
- Strong focus on performance, scalability, reliability, and data quality.
- Ability to troubleshoot complex production data issues.
- Strong attention to detail and coding discipline.
- Good communication and collaboration skills.
- Ability to work effectively with DWH, reporting, and cross-functional teams.
- Proactive approach to performance optimization and continuous improvement.
- Strong ownership of data engineering deliverables.
Interview Process
2-3 rounds of discussion.

