{"id":1315974,"url":"https://alion.io/job/digit88-senior-big-data-engineer-databricks","title":"Senior Big Data Engineer - Databricks","company":{"id":3813178,"name":"Digit88","domain":"digit88.com","url":"https://alion.io/company/digit88","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":{"grade":"C","score":66,"open_postings":14,"ghost_share":0.571,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-05T05:45:15Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agentic Workflows","optional":false},{"name":"Agile","optional":false},{"name":"Apache Kafka","optional":false},{"name":"Azure","optional":false},{"name":"Claude","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"ETL/ELT","optional":false},{"name":"pySpark","optional":false},{"name":"RAG","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Edge AI","optional":true},{"name":"HIPAA","optional":true},{"name":"Python","optional":true}],"status":"live","first_seen_at":"2026-06-09T10:53:11Z","employer_posted_date":"2026-06-09","last_verified_at":"2026-10-06T00:44:06Z","board_verified":true,"closed_at":null,"days_open":118,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":117},"description":"Sr. Data Engineer\nExperience: 8+ Years\nLocation: Bangalore - Hybrid\nType: Full-time\nAbout Digit88\nDigit88 is an AI-native product engineering partner helping startups and enterprises build, scale and operate intelligent software products.\nWe are a lean, high-impact team of 75+ technologists, backed by leaders with deep experience across startups and global enterprises. We build strong, outcome-driven engineering teams that solve complex, real-world problems.\nFrom GenAI applications and data platforms to enterprise-grade SaaS, we deliver scalable, reliable, production-ready systems. Our expertise spans AI/ML, RAG systems, agentic workflows and large-scale data engineering - enabling businesses to move from idea to production with speed and confidence.\nOur teams operate as a true extension of our clients, with full ownership and flexible engagement models focused on measurable business outcomes - not just delivery.\nWith 80+ AI implementations and proven success in scaling dedicated teams and driving significant cost efficiencies, we partner for long-term impact.\nWe bring experience across B2B and B2C SaaS, web and mobile platforms, e-commerce and domains such as Conversational AI, HealthTech, IoT, ESG/Energy and Data Engineering - thriving in fast-paced, high-ownership environments.\nThe Vision: To be the most trusted AI-native product engineering partner for innovative software companies worldwide delivering ownership, speed and measurable outcomes.\nThe Opportunity:\nAs a Senior Data Engineer, you will design, build and operate scalable data platforms and pipelines for global customers. You will work closely with customers, product teams and engineers to deliver reliable, production-grade data systems.\nYou will play a key role in shaping data architecture, data engineering best practices and AI-driven data platforms at Digit88, enabling customers to move from raw data to actionable insights and intelligent systems.\nKey Responsibilities:\nModernize legacy, non-scalable architectures and define a scalable target-state platform\nDesign and implement Medallion Architecture (Bronze, Silver, Gold) using Delta Lake and Databricks components such as Delta Live Tables (DLT), Delta Sharing and Workflows (LakeFlow, LakeBase)\nDesign and build scalable, reliable and production-ready ETL/ELT pipelines using PySpark, SQL and Databricks notebooks to ingest and transform data from diverse sources.\nCreate and manage workflows using Databricks Workflows (Jobs) or orchestration tools to automate pipelines and dependencies\nOptimize Spark jobs for performance, scalability and cost efficiency (partitioning, caching, query tuning, cluster optimization)\nImplement data quality checks (e.g., DLT Expectations) and enforce governance via Unity Catalog (access control, PII masking, lineage)\nDesign and implement event-driven and streaming pipelines (Kafka or equivalent)\nEnsure high data reliability through monitoring, observability and alerting\nRequirements:\nBE/MS in Computer Science or a related field with 8+ years of experience in data engineering\nStrong experience in ETL/ELT pipelines, data modeling, and distributed data systems, with hands-on expertise in Databricks\nDeep proficiency in PySpark, including performance optimization, job orchestration, and large-scale data processing\nGood understanding of event streaming systems such as Kafka or equivalent technologies\nExperience working with Azure data ecosystem (ADLS, Azure services, AHDS or similar data platforms)\nStrong foundation in event-driven architecture and scalable distributed systems\nProven ability to design, review, and clearly articulate system architecture\nExperience leveraging AI-assisted development tools (e.g., Claude, Antigravity, etc.) to improve productivity\nSolid experience in Agile delivery, estimation, and program execution\nExperience working with global customers (US/EU) in a client-facing role\nExcellent written and verbal communication skills across engineering, business, and customer stakeholders\nStrong analytical thinking and structured problem-solving ability\nHigh ownership, reliability, and execution focus; consistently delivers despite constraints\nStrong attention to detail while maintaining a clear big-picture perspective\nGood to have skills:\nExperience in Healthcare, EHR/EMR data migration, Clinical Trials, or Life Sciences domains (at least one)\nExposure to handling large-scale EMR/EHR integrations and reducing technical complexity across multiple data sources\nExperience working with healthcare data standards such as CDA, FHIR, and HL7, including data normalization into modern data models\nAbility to build and manage scalable data pipelines for ingestion, transformation (FHIR), data quality, governance, and near real-time processing\nUnderstanding of interoperability challenges across diverse healthcare systems and approaches to solve them\nExperience building patient-centric data platforms (e.g., Patient 360, Master Patient Index)\nFamiliarity with data privacy, security, and compliance standards such as HIPAA\nBenefits/Culture @ Digit88:\nComprehensive insurance coverage (Life, Health, Accident, parents and in-laws are optional)\nFlexible work model focused on outcomes\nAccelerated learning with non-linear growth opportunities\nFlat organization with high ownership & accountability\nOpportunity to work on cutting-edge AI and SaaS products with global customers (primarily North America, Australia, EU and UAE)\nAccomplished Global Peers - Working with some of the best engineers/professionals globally from the likes of Apple, Amazon, IBM Research, Adobe and other innovative product companies\nDirect exposure to building and scaling real-world systems across Conversational AI, Energy/Utilities, ESG, HealthTech, IoT and more\nHigh-impact roles with the ability to influence product, architecture and business outcomes globally\nLearn from a founding team of serial entrepreneurs with multiple exits - high growth, high ownership and real challenges\nThis is an exciting time to join Digit88 - build, scale and grow with us as part of our journey!","description_format":"text","description_chars":6085,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"master","optional":false},"security_clearance":false,"languages":[]},"benefits":["Flexible schedule","Growth opportunities","Insurance coverage"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","LLM & Generative AI","Influencers"],"lifecycle":[{"event":"open","at":"2026-09-26T18:35:42Z"}],"visa":[],"liveness":{"score":8,"band":"cold","label":"Long shot","p_open":1,"p_active":0.3,"p_room":0.28,"age_days":117,"expected_fill_days":40,"reasons":["conf:2","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-10-05T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/digit88-senior-big-data-engineer-databricks","json_url":"https://alion.io/job/digit88-senior-big-data-engineer-databricks.json","meta":{"generated_at":"2026-10-06T02:31:32Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3884,"day_limit":5000,"remaining_today":1116,"minute_limit":60,"resets_at":"2026-10-07T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3813178},"rest":"https://alion.io/mcp/rest/get_company?id=3813178"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fdigit88-senior-big-data-engineer-databricks"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fdigit88-senior-big-data-engineer-databricks"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fdigit88-senior-big-data-engineer-databricks"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/digit88-senior-big-data-engineer-databricks\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fdigit88-senior-big-data-engineer-databricks"}]}