{"id":1226867,"url":"https://alion.io/job/sourcebae-senior-data-engineer","title":"Senior Data Engineer","company":{"id":3776661,"name":"Sourcebae","domain":"sourcebae.com","url":"https://alion.io/company/sourcebae","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":22000,"max_usd":45000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":51},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Airflow","optional":false},{"name":"Apache Kafka","optional":false},{"name":"ClickHouse","optional":false},{"name":"Flink","optional":false},{"name":"Git","optional":false},{"name":"Java","optional":false},{"name":"Python","optional":false},{"name":"Scala","optional":false},{"name":"SQL","optional":false},{"name":"Apache NiFi","optional":true},{"name":"AWS","optional":true},{"name":"Azure","optional":true},{"name":"CI/CD","optional":true},{"name":"Dagster","optional":true},{"name":"Docker","optional":true},{"name":"GCP","optional":true},{"name":"Kubernetes","optional":true},{"name":"Prefect","optional":true},{"name":"Rest API","optional":true},{"name":"Terraform","optional":true},{"name":"WhatsApp","optional":true}],"status":"live","first_seen_at":"2026-09-22T17:37:32Z","employer_posted_date":null,"last_verified_at":"2026-09-22T17:37:32Z","board_verified":false,"closed_at":null,"days_open":6,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":6},"description":"Lead Data Engineer\nExperience: 6–8 Years\nLocation: Bengaluru / Hybrid\nEmployment Type: Full-Time\nRole Overview\nWe are looking for an experienced Lead Data Engineer with strong hands-on expertise in real-time data streaming, event processing, CDC, data integration, and modern data engineering.\nThe candidate will be responsible for designing, developing, and maintaining high-performance, production-grade data pipelines using Apache Flink, Apache Kafka, Debezium, CDC, ClickHouse, and Apache Airflow.\nThis is a hands-on engineering role requiring practical experience in building high-volume, low-latency streaming applications, developing scalable data pipelines, troubleshooting distributed systems, and optimizing data processing workloads.\nThe ideal candidate should be comfortable working across the complete data pipeline—from source systems and CDC ingestion through Kafka and Flink processing to analytical storage, APIs, dashboards, and downstream integrations.\nKey Responsibilities\n1. Real-Time Data Engineering\nDesign, develop, and maintain real-time data pipelines using Apache Flink and Apache Kafka.\nDevelop production-grade streaming applications for high-volume and low-latency workloads.\nImplement data transformation, filtering, enrichment, aggregation, and event processing.\nBuild reliable event-processing pipelines with appropriate error handling and recovery mechanisms.\nConsume and publish events across Kafka topics.\nImplement partitioning, consumer groups, offsets, and appropriate delivery mechanisms.\nTroubleshoot streaming pipeline failures, latency, throughput, and performance issues.\n2. Apache Flink\nDevelop and maintain production-grade Apache Flink jobs.\nImplement stream transformations, filtering, mapping, aggregations, joins, and windows.\nWork with event-time processing, watermarks, and state management.\nImplement Flink checkpointing, savepoints, and recovery mechanisms.\nOptimize Flink jobs for performance, scalability, and resource utilization.\nMonitor latency, throughput, failures, backpressure, and resource consumption.\nTroubleshoot state, checkpointing, backpressure, and processing issues.\n3. Apache Kafka\nDevelop Kafka-based ingestion and streaming pipelines.\nCreate and manage Kafka topics and event streams.\nWork with partitions, offsets, consumer groups, replication, and retention.\nDevelop reliable Kafka producer and consumer applications.\nHandle message ordering, retries, duplicate events, and replay scenarios.\nMonitor Kafka performance and troubleshoot consumer lag and throughput issues.\nWork with Kafka schemas and serialization formats.\n4. CDC & Debezium\nBuild CDC-based ingestion pipelines using Debezium.\nConfigure and maintain Debezium connectors.\nCapture source-system inserts, updates, and deletes.\nPublish CDC events into Kafka.\nHandle initial snapshots and incremental CDC processing.\nManage schema evolution and source-system changes.\nImplement data reconciliation and consistency checks.\nTroubleshoot CDC failures and source-to-target data issues.\n5. Data Orchestration\nDevelop and maintain data workflows using Apache Airflow or equivalent orchestration frameworks.\nBuild reusable DAGs for:\nData ingestion\nCDC workflows\nData validation\nFlink job execution\nData transformation\nClickHouse loading\nDownstream integrations\nImplement workflow dependencies, scheduling, retries, backfills, SLAs, and alerting.\nIntegrate Airflow with Kafka, Flink, Debezium, ClickHouse, APIs, and cloud services.\nMonitor workflow execution and troubleshoot failures.\nDevelop reusable operators, sensors, and workflow components where required.\nUse event-driven triggers for real-time workflows where appropriate.\n6. ClickHouse & Analytical Data\nIntegrate streaming data pipelines with ClickHouse.\nDesign efficient analytical data models.\nDevelop and optimize SQL queries.\nImplement appropriate partitioning, sorting, indexing, and retention strategies.\nOptimize data ingestion and query performance.\nSupport analytical use cases, dashboards, and reporting requirements.\n7. Data Quality & Reliability\nImplement data validation and quality checks throughout the data pipeline.\nBuild reconciliation mechanisms between source and target systems.\nMonitor data freshness, completeness, accuracy, and consistency.\nImplement error handling, retry, replay, and recovery mechanisms.\nEstablish logging and observability for critical pipelines.\nSupport incident investigation and root-cause analysis.\n8. Integration & APIs\nIntegrate streaming and analytical data with APIs, dashboards, endpoints, and downstream applications.\nDevelop data interfaces and integration components.\nWork with application teams to define data contracts and integration requirements.\nSupport future integrations and additional data consumers.\n9. Engineering Practices\nFollow modern software engineering practices including:\nGit and version control\nCode reviews\nUnit and integration testing\nCI/CD\nLogging and monitoring\nDocumentation\nDevelop reusable, scalable, and maintainable data engineering components.\nParticipate in technical design discussions and architecture reviews.\nMentor Data Engineers and contribute to engineering standards.\nRequired Skills & Experience\n6–8 years of experience in Data Engineering.\nStrong hands-on experience with Apache Flink – Mandatory/Core Requirement.\nStrong hands-on experience with Apache Kafka.\nHands-on experience with Debezium and Change Data Capture (CDC).\nStrong programming experience in Java or Scala.\nGood experience with Python is an advantage.\nStrong SQL skills.\nExperience with analytical databases; ClickHouse is highly preferred.\nHands-on experience with Apache Airflow or another data orchestration framework.\nStrong understanding of distributed systems and real-time data processing.\nExperience developing and supporting production-grade streaming pipelines.\nStrong understanding of Kafka concepts including:\nTopics\nPartitions\nOffsets\nConsumer Groups\nReplication\nRetention\nExperience with data transformation, enrichment, filtering, aggregation, and event processing.\nStrong troubleshooting and problem-solving skills for performance and reliability issues.\nFamiliarity with cloud platforms and containerized environments.\nPreferred Skills\nAdvanced Apache Flink experience, including:\nState Management\nCheckpoints\nSavepoints\nWatermarks\nEvent Time\nWindows\nBackpressure\nState Backends\nExperience with Apache Airflow, Dagster, Prefect, or Apache NiFi.\nExperience with Kafka Schema Registry.\nExperience with Avro, Protobuf, or JSON.\nExperience with Kubernetes.\nExperience with AWS, Azure, or GCP.\nExperience with CI/CD pipelines.\nExperience with Docker and containerized applications.\nExperience with Terraform or other Infrastructure as Code tools.\nExperience with data observability and monitoring tools.\nExperience building high-volume, low-latency real-time data platforms.\nExperience with REST APIs and system integrations.\nKey Technical Stack\nApache Flink | Apache Kafka | Debezium | CDC | ClickHouse | Apache Airflow | Java/Scala | Python | SQL | Kubernetes | Docker | Cloud | CI/CD | REST APIs\n\nApply Now\nInterested candidates can share their updated CV at [HIDDEN TEXT] or WhatsApp it to 8827565832.\nStay updated with our latest job opportunities and company news by following us on LinkedIn:\nSourcebae on LinkedIn","description_format":"text","description_chars":7286,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T13:06:44Z"}],"liveness":{"score":85,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.852,"p_room":1,"age_days":5,"expected_fill_days":23,"reasons":["seen:5","velocity","win:early"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/sourcebae-senior-data-engineer","json_url":"https://alion.io/job/sourcebae-senior-data-engineer.json","meta":{"generated_at":"2026-09-29T01:21:32Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":911,"day_limit":5000,"remaining_today":4089,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}