Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 11, 2026.
High Technical Impact: A unique professional challenge to shape the core data platform for a stable, market-leading global enterprise.
Competitive Compensation: Senior level base salary and annual target bonus.
Benefits Package: Cafeteria allowance and private health insurance package.
Flexibility: Hybrid working model requiring 3 days of office presence (Budapest) and offering 2 days of Home Office per week.
Responsibilities
Pipeline Engineering & Architecture: Design, build, and maintain high-performance, end-to-end data pipelines covering ingestion, processing, quality assurance, and delivery layers. Own critical data architecture decisions and technical implementations.
Lakehouse & Platform Optimization: Optimize data platform performance, reliability, and cost across Databricks Lakehouse architectures. Apply advanced tuning techniques (partitioning, clustering, materialized views, pre-aggregation) and implement medallion patterns (bronze/silver/gold with published consumer layers).
ML/AI Operations Enablement: Partner closely with data science, analytics, and ML/AI operations teams to build platforms and pipelines that enable model development, training, feature engineering, and production deployment.
Data Governance & Observability: Establish and drive best practices around data quality assurance, governance, metadata management, validation frameworks, and platform observability.
Infrastructure & CI/CD: Manage infrastructure as code and CI/CD practices across Git-based deployment workflows.
Technical Leadership & Innovation: Influence technical direction, mentor junior engineers, evaluate and adopt emerging technologies, and effectively partner across both technical and non-technical stakeholders.
Requirements
Required Qualifications
Experience: 10+ years of professional data engineering experience building production-grade data systems.
Core Stack: Deep expertise with Databricks and Apache Spark, with proven experience optimizing complex distributed workloads.
AWS & Lakehouse: Production experience with the Databricks Lakehouse on AWS (Unity Catalog, Delta Lake, Delta Live Tables, Databricks SQL) alongside core supporting AWS services (S3, IAM, Lambda, RDS).
Languages & Querying: Expert SQL skills and strong programming proficiency in Python (PySpark). (Scala knowledge is a plus).
Architectural Mastery: Solid understanding of Lakehouse and data warehouse architectures, medallion patterns, and dimensional modeling.
Data Domain Background: Professional experience at companies where data ingestion, delivery, and consumption are core business functions.
Collaboration & Leadership: Strong communication skills with a proven ability to mentor junior engineers, influence technical direction, and partner across diverse stakeholders.
Preferred Qualifications
Hands-on experience with data science workflows (feature engineering, model input preparation, evaluation support) and production ML/AI operations (ML monitoring, retraining pipelines, model deployment).
Exposure to LLM/generative AI infrastructure, RAG systems, or embeddings pipelines in production.
Infrastructure-as-Code and CI/CD proficiency (Databricks Asset Bundles, Terraform, Git-based workflows).
Advanced knowledge of open table formats (Delta Lake, Apache Iceberg), data quality frameworks, metadata management, and observability/monitoring tools.
Cloud cost optimization expertise and/or contributions to open-source data engineering projects.
Technical Skills & Domain Knowledge
Distributed computing fundamentals, large-scale data processing optimization, and performance benchmarking/profiling.
Real-time streaming, event-driven architectures, API design, and integration patterns for data consumption.
Security, encryption, and compliance requirements for sensitive enterprise data.
Familiarity with AI-assisted development tools (Claude Code, Cursor, Databricks Genie).

