812,549open jobs
52,293companies
130,976added this week
Browse all
Salary
$33k – $37k per year
Location
In office (Noida)
Seniority
Principal · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 10, 2026.

Overview
Company
Impact
Profile match
Algoworks is an artificial intelligence, engineering services and experience transformation firm that designs, builds and runs software, data and CRM platforms for enterprise clients. The company has offices across the United States, Europe, South America and India, has worked with Fortune 500 organizations for more than 20 years, and runs its main delivery centre in Noida with further teams in Gurugram and Hyderabad. Its current openings are mostly in Noida and cover ServiceNow, Salesforce and Microsoft Dynamics specialists, data and backend engineers in Python and Java, DevOps and mobile QA engineers, solution architects and delivery managers.

Role: Principal Data Engineer - Real-time Data Platform

Location: India, Remote

Experience: 10+ years

Algoworks

www.algoworks.com

About the company

Algoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.

For over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.

At Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.

Through collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.

Follow the video below to know about us! Clipchamp

Role overview

We are looking for a hands-on Lead Data Engineer to provide technical leadership for high-volume data ingestion and processing, with a strong focus on real-time CDC, Databricks, SQL Server, Debezium and Azure Event Hubs.

This role will own the technical architecture and engineering direction across Full Load / Batch and Real-Time CDC pipelines, with real-time streaming expected to become the primary long-term ingestion pattern.

The ideal candidate is a strong technical leader and architect who remains hands-on, can guide and mentor engineers and can design production-grade ingestion solutions for scalability, resiliency, performance and data integrity.

Key responsibilities:

  • Lead the architecture and technical implementation of batch, full-load, incremental and real-time CDC pipelines.
  • Design high-volume ingestion from SQL Server using CDC and Debezium.
  • Build scalable event-driven pipelines using Azure Event Hubs and Databricks.
  • Design and optimize Databricks pipelines for large-scale data ingestion and transformation.
  • Implement robust error handling, retry, replay, checkpointing, recovery and idempotency.
  • Design solutions for schema drift and schema evolution without disrupting downstream processing.
  • Design and optimize Delta Lake / Delta Tables, including partitioning, compaction, data layout and performance optimization.
  • Optimize pipeline throughput, latency, parallelism, resource utilization and processing windows.
  • Establish monitoring and observability for CDC lag, connector health, consumer lag, pipeline failures, throughput and processing latency.
  • Implement reconciliation and data-quality controls to ensure source-to-target completeness and accuracy.
  • Provide technical direction, perform design/code reviews, mentor engineers and establish engineering best practices.
  • Drive technical readiness for scaling ingestion across significantly more clients, databases, tables and data volumes.

Required skills and qualifications:

  • Bachelor’s or master's degree in computer science, Information Technology, Business, or related field (or equivalent practical experience).
  • 10+ years of Data Engineering / Software Engineering experience.
  • Strong hands-on experience with Databricks and Delta Lake.
  • Strong experience designing and operating Databricks data pipelines at scale.
  • Deep understanding of:
    • Pipeline design and orchestration
    • Error handling and recovery
    • Schema drift
    • Schema evolution
    • Idempotent data processing
    • Delta Tables
    • Data partitioning and optimization
    • Performance tuning
  • Strong hands-on experience with SQL Server CDC, transaction logs, LSNs and high-volume transactional databases.
  • Experience with Debezium SQL Server Connector, including configuration, offsets, snapshots, recovery and schema changes.
  • Strong experience with Azure Event Hubs, including partitioning, consumer groups, scaling, throughput and checkpointing.
  • Deep understanding of batch, micro-batch, streaming and event-driven data architectures.
  • Strong experience with Python/PySpark, SQL, Azure Data Lake and distributed data processing.
  • Experience designing production-grade solutions for retry, replay, fault tolerance, duplicate handling, reconciliation and observability.
  • Strong performance engineering and troubleshooting skills across large-scale data pipelines.
  • Ability to provide technical leadership, architecture guidance, mentoring and hands-on engineering support.

Must have skills:

  • 10+ years of Data Engineering / Software Engineering experience.
  • 3+ years working with production-scale CDC or real-time streaming architectures.
  • Strong production experience with Databricks and Delta Lake.
  • Experience processing millions to billions of records.
  • Experience with multi-client or multi-tenant ingestion architectures.
  • Experience implementing Medallion / Bronze-Silver-Gold architectures.

Good to have skills:

  • Experience with Apache Kafka / Kafka Connect and streaming ecosystems.
  • Knowledge of Azure Data Factory, Azure Functions and Azure Monitor.
  • Experience with Infrastructure as Code (Terraform/ARM/Bicep) and CI/CD for data platforms.
  • Familiarity with Unity Catalog, Databricks Workflows and advanced Spark optimization.

Key success criteria:

The person in this role should be able to:

  • Establish a scalable architecture for batch and real-time ingestion.
  • Scale pipelines across substantially more databases, clients, tables and data volumes.
  • Improve Databricks pipeline performance and processing windows.
  • Deliver reliable high-volume CDC without sustained lag, duplication, or data loss.
  • Handle schema changes and schema drift without destabilizing ingestion.
  • Ensure pipelines are idempotent and safely recoverable/replayable following failures.
  • Optimize Delta Tables and downstream processing for performance and scalability.
  • Provide clear technical leadership and mentoring for the ingestion engineering team.

Interview process

2 rounds of discussion.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
812,549 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Noida
≈ $32k – $73k per year (Estimated) • Hybrid • Contractor • 8+ years exp • Hyderabad
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Apache Kafka
AI/ML
Spark
DevOps
Azure DevOps
Azure
CI/CD
Git
AWS
Analytics
Tableau
Power BI
ETL/ELT
Azure Data Factory
AWS Glue
Dimensional Modeling
Apply
Data Architect 13 days ago
≈ $33k – $77k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Pune
Python
SQL
ABAP
Python
pySpark
ABAP
CDS Views
Databases
Databricks
Microsoft Fabric
SAP BW
AI/ML
Spark
DevOps
Azure
Analytics
Power BI
ETL/ELT
Azure Data Factory
Collibra
Management
Agile
Apply
≈ $36k – $78k per year (Estimated) • Remote (India) • Full-Time • Bengaluru • Hyderabad • Pune • Chennai • Greater Noida
DevOps
Incident Management
SLI/SLO/SLA
Cybersecurity
Least Privilege
Analytics
Power BI
Apply
Staff Data Architect 8 hours ago
≈ $32k – $75k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
SQL
Databases
Google BigQuery
BigQuery
DevOps
GCP
CI/CD
AWS
Incident Management
Analytics
ETL/ELT
Collibra
Apply
≈ $40k – $93k per year (Estimated) • Remote (India) • 9+ years exp
Python
SQL
Databases
Snowflake
AI/ML
Claude
dbt
Apply
≈ $93k – $197k per year (Estimated) • In office • 5+ years exp
Python
SQL
Scala
Databases
Databricks
MS SQL
Apache Kafka
Microsoft Fabric
AI/ML
Spark
DevOps
Azure
Analytics
Power BI
Dimensional Modeling
Apply
≈ $32k – $80k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Athens
Python
SQL
Databases
Databricks
AI/ML
Copilot
Claude
Dagster
dbt
DevOps
Azure
AWS
Analytics
ETL/ELT
Azure Data Factory
Apply
$155k – $163k per year • In office • 4+ years exp • Bachelor's Degree
Python
SQL
Analytics
ETL/ELT
Management
Agile
Apply
≈ $66k – $140k per year (Estimated) • Remote (likely United States) • 10+ years exp • Bachelor's Degree
SQL
DevOps
Azure DevOps
Azure
Analytics
Power BI
Management
Power Automate
Power Apps
Apply
≈ $103k – $207k per year (Estimated) • Remote (likely United States) • 4+ years exp • Bachelor's Degree
SQL
DevOps
Azure DevOps
Azure
Analytics
Power BI
Management
Power Automate
Apply
Lead Data Engineer 2 days ago
$34k – $42k per year • In office • Full-Time • 10+ years exp • Noida
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Microsoft Fabric
AI/ML
Spark
DevOps
Azure
CI/CD
Git
Analytics
ETL/ELT
Azure Data Factory
Management
Agile
Apply
$31k – $39k per year • In office • Full-Time • 10+ years exp • Noida
Python
SQL
Databases
RabbitMQ
DevOps
Rest API
CI/CD
AWS
Apply
Data Engineer 1 month ago
$31k – $37k per year • In office • Full-Time • 6+ years exp • Noida
Python
SQL
Python
pySpark
Databases
Snowflake
AI/ML
Spark
Dagster
dbt
Frontend
GraphQL
DevOps
CI/CD
Apply
$10k – $31k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Apache Kafka
AI/ML
Spark
DevOps
Terraform
Azure
CI/CD
Bicep
Analytics
Azure Data Factory
Apply
Senior Engineer 1 day ago
$19k – $26k per year • In office • Full-Time • 5+ years exp • Noida
Apex
DevOps
Datadog
CI/CD
Incident Management
SLI/SLO/SLA
Management
ITSM
Apply
≈ $28k – $59k per year (Estimated) • In office • Full-Time • Noida
AI/ML
Machine Learning
Cybersecurity
ISO 27001
SOC 2
GDPR
HIPAA
Apply
≈ $15k – $31k per year (Estimated) • Remote (India) • Full-Time • 10+ years exp • Bachelor's Degree • Noida
Python
Java
SQL
Scala
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
AI/ML
Spark
AI Agents
Flink
Edge AI
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
≈ $39k – $88k per year (Estimated) • In office • 10+ years exp • Noida • Bengaluru • Chennai • Thiruvananthapuram • Kochi
Python
SQL
Python
pySpark
Databases
Azure SQL Database
Microsoft Fabric
AI/ML
Spark
dbt
DevOps
Azure
Analytics
Power BI
Azure Data Factory
Apply
≈ $25k – $62k per year (Estimated) • In office • Noida
Python
Go
Java
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Cybersecurity
GDPR
HIPAA
Apply
In office • Full-Time • Master's Degree • Noida
Python
C++
AI/ML
Prompt Engineering
NLP
LLM
RAG
Semantic Search
OpenAI
Anthropic
Semantic Search
DevOps
CI/CD
Git
Docker
Kubernetes
Management
Agile
Apply
See all jobs
This is one of many
812,549 more open roles from verified company boards, updated every day.