{"id":1101015,"url":"https://alion.io/job/stermedia-senior-data-engineer-databricks-expert","title":"Senior Data Engineer / Databricks Expert","company":{"id":28404,"name":"Stermedia","domain":"stermedia.ai","url":"https://alion.io/company/stermedia","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Recruitee","truth_index":{"grade":"B","score":75,"open_postings":5,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-28T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Wrocław, Poland"],"countries":["PL"],"hiring_countries":["PL"],"hiring_countries_total":1,"salary":{"min":20000,"max":30000,"currency":"PLN","period":"month","gross":null,"usd_annual":93900},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"ETL/ELT","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"Machine Learning","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2021-06-18T08:02:07Z","employer_posted_date":"2021-06-18","last_verified_at":"2026-09-28T15:10:24Z","board_verified":true,"closed_at":null,"days_open":1928,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":1927},"description":"Stermedia.ai is looking for an experienced Senior Data Engineer with strong Databricks expertise to join a data platform team building a global analytics environment for a leading pharmaceutical organization.\nThe platform integrates data from more than 80 facilities worldwide, including laboratories, manufacturing systems, logistics platforms, and operational environments. Its goal is to transform distributed datasets into reliable, analytics-ready data that supports reporting, operational monitoring, and data-driven decision-making.\nYou’ll contribute to the full delivery lifecycle - from direct client discussions and business requirements analysis to solution design, implementation, and production deployment. Working alongside international data engineering, analytics, and machine learning teams, you’ll focus on building scalable data pipelines and transformation layers using Databricks, Apache Spark, Python, SQL, and Delta Lake.\nAbout the project\nThe project focuses on building a scalable data platform for operational data from pharmaceutical laboratories, production systems, and global supply chains.\nData from multiple regions is consolidated into a centralized analytics environment. The role focuses on using Databricks and a lakehouse architecture to ingest, process, and organize this data into reliable datasets for downstream analytics.\nAs a Databricks expert, you’ll design and maintain data pipelines, optimize distributed processing, and support data quality, governance, and production reliability. DBT is a complementary skill for SQL-based transformations and analytics engineering workflows.\nRequirements\nDesign, develop, and maintain data pipelines in Databricks using Python, PySpark, and SQL\n\nBuild and maintain Delta Lake tables and transformation layers following bronze, silver, and gold architecture\n\nIntegrate data from multiple source systems into consistent, analytics-ready datasets\n\nImplement incremental processing, change data capture, and schema evolution where required\n\nDevelop and orchestrate production workflows, including dependencies, scheduling, retries, and monitoring\n\nOptimize Spark workloads, SQL queries, and compute usage to improve performance and manage costs\n\nImplement automated data quality checks, validation, and pipeline tests\n\nSupport data access management, governance, and lineage using Unity Catalog\n\nMaintain reusable code, technical documentation, and Git-based development workflows\n\nCollaborate with business stakeholders, data engineers, analysts, and machine learning specialists to translate requirements into reliable data solutions\n\nRequired skills\nStrong commercial experience with Databricks, including developing and operating production data pipelines\n\nAdvanced knowledge of Python, PySpark, and SQL\n\nHands-on experience with Apache Spark, distributed data processing, and performance troubleshooting\n\nPractical experience with Delta Lake, including incremental loads, merge operations, and schema management\n\nStrong understanding of lakehouse architecture, ETL/ELT patterns, and data modeling\n\nExperience orchestrating and monitoring workflows in Databricks\n\nPractical knowledge of Unity Catalog, including permissions, data organization, and lineage\n\nExperience optimizing pipeline performance and compute resource usage\n\nFamiliarity with cloud storage and services in at least one major cloud environment: Azure, AWS, or GCP\n\nExperience with Git, code reviews, automated testing, and CI/CD workflows\n\nStrong analytical and communication skills, with the ability to work directly with international stakeholders\n\nGood command of English\n\nNice to have\nExperience with DBT, including models, tests, macros, and integration with Databricks\n\nExperience with Structured Streaming and Databricks Auto Loader\n\nFamiliarity with Terraform and infrastructure as code\n\nExperience supporting datasets and workflows used by machine learning teams\n\nExperience building global analytics platforms or integrating data from multiple facilities\n\nFamiliarity with pharmaceutical, manufacturing, laboratory, or supply chain data\n\nDatabricks certifications in data engineering\n\nTechnology focus\nData platform: Databricks\n\nProcessing: Apache Spark, PySpark\n\nLanguages: Python, SQL\n\nStorage and table format: Delta Lake, cloud object storage\n\nArchitecture: Lakehouse, bronze / silver / gold layers\n\nGovernance: Unity Catalog\n\nOrchestration: Databricks jobs and workflows\n\nDevelopment and delivery: Git, CI/CD, automated testing\n\nComplementary analytics tooling: DBT\n\nWe offer you\nOpportunities to work with modern data engineering and machine learning technologies\n\nAn annual self-development budget\n\nThe opportunity to contribute to a variety of interesting projects\n\nInternal workshops and knowledge-sharing sessions\n\nSupport for personal branding through articles, conference talks, and leading internal workshops\n\nFlexible working hours\n\nThe possibility of remote work\n\nA chillout room, free beverages, and team and company events\n\nA friendly atmosphere\n\nMultiSport\n\nLuxMed\n\nSalary\n20,000-30,000 PLN + VAT (B2B)","description_format":"text","description_chars":5082,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":["Company events","Flexible schedule"],"hiring_locations":[{"name":"Poland","iso":"PL","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Mobile App Development Services","Custom Software Development","AI Consulting & Integration"],"lifecycle":[{"event":"open","at":"2026-09-21T22:38:46Z"}],"liveness":{"score":3,"band":"cold","label":"Long shot","p_open":1,"p_active":0.096,"p_room":0.28,"age_days":1927,"expected_fill_days":30,"reasons":["conf:11","win:tail","crowd:"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":93900,"is_top_pay":false},"html_url":"https://alion.io/job/stermedia-senior-data-engineer-databricks-expert","json_url":"https://alion.io/job/stermedia-senior-data-engineer-databricks-expert.json","meta":{"generated_at":"2026-09-28T23:33:02Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1223,"day_limit":5000,"remaining_today":3777,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}