{"id":1247485,"url":"https://alion.io/job/algoworks-principal-data-engineer-real-time-data-platform-2","title":"Principal Data Engineer – Real-time Data Platform","company":{"id":2091826,"name":"Algoworks","domain":"algoworks.com","url":"https://alion.io/company/algoworks-com","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Noida, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":3200000,"max":3500000,"currency":"INR","period":"year","gross":null,"usd_annual":36687},"salary_estimate":null,"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"SQL","optional":false},{"name":"Apache Kafka","optional":true},{"name":"Azure Data Factory","optional":true},{"name":"Bicep","optional":true},{"name":"CI/CD","optional":true},{"name":"Spark","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-10T05:50:39Z","employer_posted_date":"2026-09-10","last_verified_at":"2026-09-26T23:35:33Z","board_verified":true,"closed_at":null,"days_open":16,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":16},"description":"Role: Principal Data Engineer - Real-time Data Platform\nLocation: India, Remote\nExperience: 10+ years\nAlgoworks\nwww.algoworks.com\nAbout the company\nAlgoworks is an award-winning artificial intelligence, engineering services and experience transformation firm with offices across the United States, Europe, South America and India. We bring together a global team of engineers, architects, designers, researchers and operators united by rigor, accountability and a commitment to delivering measurable results.\nFor over 20 years, Algoworks has partnered with Fortune 500 organizations across the Americas, Europe and Asia to define, build and run technology that drives meaningful business outcomes. Our work combines human-centered design, engineering excellence and AI-powered capabilities to solve complex challenges with clarity and precision. Innovation, particularly in the responsible application of AI, is embedded in how teams approach problem-solving and continuous improvement.\nAt Algoworks, growth is continuous and closely tied to impact. Teams collaborate across geographies and disciplines, strengthening outcomes through shared insight and collective expertise. The culture values transparency, open dialogue and an environment where every voice is heard and contribution is recognized.\nThrough collaboration, accountability and a focus on results, Algoworks operates at the intersection of technology and people, building not only advanced systems but strong global teams that elevate performance and create lasting impact.\nFollow the video below to know about us! Clipchamp\nRole overview\nWe are looking for a hands-on Lead Data Engineer to provide technical leadership for high-volume data ingestion and processing, with a strong focus on real-time CDC, Databricks, SQL Server, Debezium and Azure Event Hubs.\nThis role will own the technical architecture and engineering direction across Full Load / Batch and Real-Time CDC pipelines, with real-time streaming expected to become the primary long-term ingestion pattern.\nThe ideal candidate is a strong technical leader and architect who remains hands-on, can guide and mentor engineers and can design production-grade ingestion solutions for scalability, resiliency, performance and data integrity.\nKey responsibilities:\nLead the architecture and technical implementation of batch, full-load, incremental and real-time CDC pipelines.\nDesign high-volume ingestion from SQL Server using CDC and Debezium.\nBuild scalable event-driven pipelines using Azure Event Hubs and Databricks.\nDesign and optimize Databricks pipelines for large-scale data ingestion and transformation.\nImplement robust error handling, retry, replay, checkpointing, recovery and idempotency.\nDesign solutions for schema drift and schema evolution without disrupting downstream processing.\nDesign and optimize Delta Lake / Delta Tables, including partitioning, compaction, data layout and performance optimization.\nOptimize pipeline throughput, latency, parallelism, resource utilization and processing windows.\nEstablish monitoring and observability for CDC lag, connector health, consumer lag, pipeline failures, throughput and processing latency.\nImplement reconciliation and data-quality controls to ensure source-to-target completeness and accuracy.\nProvide technical direction, perform design/code reviews, mentor engineers and establish engineering best practices.\nDrive technical readiness for scaling ingestion across significantly more clients, databases, tables and data volumes.\nRequired skills and qualifications:\nBachelor’s or master's degree in computer science, Information Technology, Business, or related field (or equivalent practical experience).\n10+ years of Data Engineering / Software Engineering experience.\nStrong hands-on experience with Databricks and Delta Lake.\nStrong experience designing and operating Databricks data pipelines at scale.\nDeep understanding of:Pipeline design and orchestration\nError handling and recovery\nSchema drift\nSchema evolution\nIdempotent data processing\nDelta Tables\nData partitioning and optimization\nPerformance tuning\n\nStrong hands-on experience with SQL Server CDC, transaction logs, LSNs and high-volume transactional databases.\nExperience with Debezium SQL Server Connector, including configuration, offsets, snapshots, recovery and schema changes.\nStrong experience with Azure Event Hubs, including partitioning, consumer groups, scaling, throughput and checkpointing.\nDeep understanding of batch, micro-batch, streaming and event-driven data architectures.\nStrong experience with Python/PySpark, SQL, Azure Data Lake and distributed data processing.\nExperience designing production-grade solutions for retry, replay, fault tolerance, duplicate handling, reconciliation and observability.\nStrong performance engineering and troubleshooting skills across large-scale data pipelines.\nAbility to provide technical leadership, architecture guidance, mentoring and hands-on engineering support.\nMust have skills:\n10+ years of Data Engineering / Software Engineering experience.\n3+ years working with production-scale CDC or real-time streaming architectures.\nStrong production experience with Databricks and Delta Lake.\nExperience processing millions to billions of records.\nExperience with multi-client or multi-tenant ingestion architectures.\nExperience implementing Medallion / Bronze-Silver-Gold architectures.\nGood to have skills:\nExperience with Apache Kafka / Kafka Connect and streaming ecosystems.\nKnowledge of Azure Data Factory, Azure Functions and Azure Monitor.\nExperience with Infrastructure as Code (Terraform/ARM/Bicep) and CI/CD for data platforms.\nFamiliarity with Unity Catalog, Databricks Workflows and advanced Spark optimization.\nKey success criteria:\nThe person in this role should be able to:\nEstablish a scalable architecture for batch and real-time ingestion.\nScale pipelines across substantially more databases, clients, tables and data volumes.\nImprove Databricks pipeline performance and processing windows.\nDeliver reliable high-volume CDC without sustained lag, duplication, or data loss.\nHandle schema changes and schema drift without destabilizing ingestion.\nEnsure pipelines are idempotent and safely recoverable/replayable following failures.\nOptimize Delta Tables and downstream processing for performance and scalability.\nProvide clear technical leadership and mentoring for the ingestion engineering team.\nInterview process\n2 rounds of discussion.","description_format":"text","description_chars":6477,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Science & Engineering","Mobile Games","Engineering Services"],"lifecycle":[{"event":"open","at":"2026-09-25T17:58:41Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.695,"p_room":0.9,"age_days":15,"expected_fill_days":24,"reasons":["conf:4","win:mid"],"computed_at":"2026-09-26T05:45:00Z"},"pay":{"stated_usd_annual":36687,"is_top_pay":false},"html_url":"https://alion.io/job/algoworks-principal-data-engineer-real-time-data-platform-2","json_url":"https://alion.io/job/algoworks-principal-data-engineer-real-time-data-platform-2.json","meta":{"generated_at":"2026-09-27T05:20:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4680,"day_limit":5000,"remaining_today":320,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}