925,716open jobs
56,515companies
155,668added this week
Browse all
Salary
≈ $29k – $51k per year (Estimated)
Location
Hybrid (Bengaluru, India)
Seniority
Staff · 6+ years exp
Employment
Full-Time

First seen by Alion on Sep 28, 2026.

Overview
Company
Impact
Profile match
Meet Sourcebae. We’re an AI-powered staffing company that helps startups and enterprises hire faster with vetted talent, a simple process, and real support.

Lead Data Engineer

Experience: 6–8 Years

Location: Bengaluru / Hybrid

Employment Type: Full-Time

Role Overview

We are looking for an experienced Lead Data Engineer with strong hands-on expertise in real-time data streaming, event processing, CDC, data integration, and modern data engineering.

The candidate will be responsible for designing, developing, and maintaining high-performance, production-grade data pipelines using Apache Flink, Apache Kafka, Debezium, CDC, ClickHouse, and Apache Airflow.

This is a hands-on engineering role requiring practical experience in building high-volume, low-latency streaming applications, developing scalable data pipelines, troubleshooting distributed systems, and optimizing data processing workloads.

The ideal candidate should be comfortable working across the complete data pipeline—from source systems and CDC ingestion through Kafka and Flink processing to analytical storage, APIs, dashboards, and downstream integrations.

Key Responsibilities

1. Real-Time Data Engineering

  • Design, develop, and maintain real-time data pipelines using Apache Flink and Apache Kafka.
  • Develop production-grade streaming applications for high-volume and low-latency workloads.
  • Implement data transformation, filtering, enrichment, aggregation, and event processing.
  • Build reliable event-processing pipelines with appropriate error handling and recovery mechanisms.
  • Consume and publish events across Kafka topics.
  • Implement partitioning, consumer groups, offsets, and appropriate delivery mechanisms.
  • Troubleshoot streaming pipeline failures, latency, throughput, and performance issues.

2. Apache Flink

  • Develop and maintain production-grade Apache Flink jobs.
  • Implement stream transformations, filtering, mapping, aggregations, joins, and windows.
  • Work with event-time processing, watermarks, and state management.
  • Implement Flink checkpointing, savepoints, and recovery mechanisms.
  • Optimize Flink jobs for performance, scalability, and resource utilization.
  • Monitor latency, throughput, failures, backpressure, and resource consumption.
  • Troubleshoot state, checkpointing, backpressure, and processing issues.

3. Apache Kafka

  • Develop Kafka-based ingestion and streaming pipelines.
  • Create and manage Kafka topics and event streams.
  • Work with partitions, offsets, consumer groups, replication, and retention.
  • Develop reliable Kafka producer and consumer applications.
  • Handle message ordering, retries, duplicate events, and replay scenarios.
  • Monitor Kafka performance and troubleshoot consumer lag and throughput issues.
  • Work with Kafka schemas and serialization formats.

4. CDC & Debezium

  • Build CDC-based ingestion pipelines using Debezium.
  • Configure and maintain Debezium connectors.
  • Capture source-system inserts, updates, and deletes.
  • Publish CDC events into Kafka.
  • Handle initial snapshots and incremental CDC processing.
  • Manage schema evolution and source-system changes.
  • Implement data reconciliation and consistency checks.
  • Troubleshoot CDC failures and source-to-target data issues.

5. Data Orchestration

  • Develop and maintain data workflows using Apache Airflow or equivalent orchestration frameworks.
  • Build reusable DAGs for:
  • Data ingestion
  • CDC workflows
  • Data validation
  • Flink job execution
  • Data transformation
  • ClickHouse loading
  • Downstream integrations
  • Implement workflow dependencies, scheduling, retries, backfills, SLAs, and alerting.
  • Integrate Airflow with Kafka, Flink, Debezium, ClickHouse, APIs, and cloud services.
  • Monitor workflow execution and troubleshoot failures.
  • Develop reusable operators, sensors, and workflow components where required.
  • Use event-driven triggers for real-time workflows where appropriate.

6. ClickHouse & Analytical Data

  • Integrate streaming data pipelines with ClickHouse.
  • Design efficient analytical data models.
  • Develop and optimize SQL queries.
  • Implement appropriate partitioning, sorting, indexing, and retention strategies.
  • Optimize data ingestion and query performance.
  • Support analytical use cases, dashboards, and reporting requirements.

7. Data Quality & Reliability

  • Implement data validation and quality checks throughout the data pipeline.
  • Build reconciliation mechanisms between source and target systems.
  • Monitor data freshness, completeness, accuracy, and consistency.
  • Implement error handling, retry, replay, and recovery mechanisms.
  • Establish logging and observability for critical pipelines.
  • Support incident investigation and root-cause analysis.

8. Integration & APIs

  • Integrate streaming and analytical data with APIs, dashboards, endpoints, and downstream applications.
  • Develop data interfaces and integration components.
  • Work with application teams to define data contracts and integration requirements.
  • Support future integrations and additional data consumers.

9. Engineering Practices

  • Follow modern software engineering practices including:
  • Git and version control
  • Code reviews
  • Unit and integration testing
  • CI/CD
  • Logging and monitoring
  • Documentation
  • Develop reusable, scalable, and maintainable data engineering components.
  • Participate in technical design discussions and architecture reviews.
  • Mentor Data Engineers and contribute to engineering standards.

Required Skills & Experience

  • 6–8 years of experience in Data Engineering.
  • Strong hands-on experience with Apache Flink – Mandatory/Core Requirement.
  • Strong hands-on experience with Apache Kafka.
  • Hands-on experience with Debezium and Change Data Capture (CDC).
  • Strong programming experience in Java or Scala.
  • Good experience with Python is an advantage.
  • Strong SQL skills.
  • Experience with analytical databases; ClickHouse is highly preferred.
  • Hands-on experience with Apache Airflow or another data orchestration framework.
  • Strong understanding of distributed systems and real-time data processing.
  • Experience developing and supporting production-grade streaming pipelines.
  • Strong understanding of Kafka concepts including:
  • Topics
  • Partitions
  • Offsets
  • Consumer Groups
  • Replication
  • Retention
  • Experience with data transformation, enrichment, filtering, aggregation, and event processing.
  • Strong troubleshooting and problem-solving skills for performance and reliability issues.
  • Familiarity with cloud platforms and containerized environments.

Preferred Skills

  • Advanced Apache Flink experience, including:
  • State Management
  • Checkpoints
  • Savepoints
  • Watermarks
  • Event Time
  • Windows
  • Backpressure
  • State Backends
  • Experience with Apache Airflow, Dagster, Prefect, or Apache NiFi.
  • Experience with Kafka Schema Registry.
  • Experience with Avro, Protobuf, or JSON.
  • Experience with Kubernetes.
  • Experience with AWS, Azure, or GCP.
  • Experience with CI/CD pipelines.
  • Experience with Docker and containerized applications.
  • Experience with Terraform or other Infrastructure as Code tools.
  • Experience with data observability and monitoring tools.
  • Experience building high-volume, low-latency real-time data platforms.
  • Experience with REST APIs and system integrations.

Key Technical Stack

Apache Flink | Apache Kafka | Debezium | CDC | ClickHouse | Apache Airflow | Java/Scala | Python | SQL | Kubernetes | Docker | Cloud | CI/CD | REST APIs

Apply Now

Interested candidates can share their updated CV at [HIDDEN TEXT] or WhatsApp it to 8827565832.

Stay updated with our latest job opportunities and company news by following us on LinkedIn:

Sourcebae on LinkedIn

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
925,716 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Bengaluru
Data Engineer 1 day ago
≈ $72k – $159k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Kansas City
SQL
Databases
Snowflake
Db2
Delta Lake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Copilot
TensorFlow
DevOps
Azure
CI/CD
AWS
Analytics
Power BI
ETL/ELT
Dimensional Modeling
Management
Agile
Apply
Senior Data Engineer 2 months ago
≈ $67k – $100k per year (Estimated) • Remote (Poland)
Python
SQL
Databases
Snowflake
AI/ML
dbt
OpenAI
Analytics
ETL/ELT
Apply
≈ $86k – $165k per year (Estimated) • In office
SQL
Databases
SAP BW
Analytics
ETL/ELT
Data Vault
Dimensional Modeling
Master Data Management
Apply
$78k – $147k per year • Hybrid • Full-Time • Bachelor's Degree • Washington
Python
SQL
Databases
Apache Iceberg
Amazon Aurora
Trino
AI/ML
AWS Bedrock
Machine Learning
DevOps
AWS
Amazon S3
Analytics
AWS Glue
Apply
≈ $108k – $210k per year (Estimated) • Hybrid • Full-Time • London
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Spark
dbt
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Analytics
Tableau
Power BI
ETL/ELT
Azure Data Factory
Apply
≈ $70k – $132k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Everett
Python
SQL
Databases
Snowflake
Oracle
DevOps
Rest API
CI/CD
Git
Analytics
Power BI
Apply
≈ $19k – $48k per year (Estimated) • In office • 5+ years exp • Master's Degree • Cape Town
Python
SQL
AI/ML
Time Series Forecasting
Analytics
Tableau
Power BI
Management
Agile
Apply
Senior Data Developer 9 hours ago
$86k – $106k per year • Hybrid • 8+ years exp • Bachelor's Degree
Python
SQL
Databases
PostgreSQL
Databricks
MS SQL
Apache Kafka
AI/ML
Spark
DevOps
Prometheus
Azure
CI/CD
Grafana
Analytics
ETL/ELT
Apply
Hybrid • 3+ years exp • Bachelor's Degree
Python
SQL
Analytics
Collibra
Apply
$38k – $40k per year • In office • 3+ years exp • Bachelor's Degree • Mozambique
Python
JavaScript
SQL
Ruby
Databases
MySQL
PostgreSQL
Analytics
Power BI
Metabase
Superset
Apply
≈ $22k – $45k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bengaluru
Python
Java
SQL
Scala
Databases
ClickHouse
Apache Kafka
AI/ML
Airflow
Dagster
Prefect
Flink
DevOps
Rest API
Terraform
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
Apache NiFi
Management
WhatsApp
Apply
≈ $21k – $44k per year (Estimated) • In office • 5+ years exp • India
Databases
PostgreSQL
ClickHouse
Milvus
DevOps
ZooKeeper
GitHub Actions
etcd
CI/CD
ArgoCD
Jenkins
Kubernetes
GitHub
Linux
Apply
$10k – $25k per year (gross) • In office • 3+ years exp • India
JavaScript
SQL
C#
C#
.NET
DevOps
Rest API
Azure
Windows
SOAP
Analytics
Power BI
Management
Power Automate
Power Apps
Outlook
SharePoint
Apply
$10k – $25k per year (gross) • In office • 5+ years exp • India
Apex
Apex
Lightning Web Components
Salesforce Industries
Apply
Full Stack Engineer 4 days ago
≈ $16k – $40k per year (Estimated) • In office • 8+ years exp • Hyderabad
Python
JavaScript
TypeScript
SQL
Node JS
Python
FastAPI
Node JS
Axios
Databases
MySQL
PostgreSQL
Redis
Frontend
Redux
Next.js
React.js
Mobile
Clean Architecture
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
Amazon S3
Amazon ECS
QA
Jest
Pytest
Apply
Sr Workday Analyst 1 day ago
≈ $13k – $26k per year (Estimated) • In office • Full-Time • 3+ years exp • Chennai • Coimbatore • Hyderabad • Mumbai • Pune
Analytics
Microsoft Excel
Apply
≈ $31k – $69k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Bengaluru
Apply
≈ $21k – $48k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Ruby
AI/ML
AI Agents
DevOps
Jenkins
DHCP
Wi-Fi
Management
Jira
Apply
≈ $53k – $110k per year (Estimated) • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Pune • Bengaluru • Mumbai • Hyderabad
AI/ML
Edge AI
DevOps
AIOps
Management
Agile
Apply
≈ $18k – $40k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru
Apply
See all jobs
This is one of many
925,716 more open roles from verified company boards, updated every day.