We are looking for an experienced Data Engineer to join a long-term project in Warsaw, focusing on data processing and analytics. You will be part of a dynamic team working with Microsoft Fabric to design and implement data pipelines, ensuring high-quality data integration and transformation processes. This hybrid role offers the flexibility of remote work while collaborating closely with business teams and stakeholders. If you are passionate about leveraging data to drive insights and solutions, this opportunity may be the right fit for you.
Your tasks
Designing and implementing data pipelines in Data Factory and PySpark notebooks in Microsoft Fabric
Onboarding file and database sources, including technical validation and format standardization
Transforming and preparing batch data for analysis
Implementing data quality validation rules and error reporting mechanisms
Versioning datasets and recording metadata of the process (parameters, sources, versions)
Working in a Git environment, managing code repositories, CI/CD deployments, and conducting code reviews
Collaborating with the Risk Team to establish transformation rules
Ensuring compliance with data privacy regulations and integrating company security standards into the Azure environment
Optimizing processes for performance and cost efficiency
Requirements
Minimum 3 years of experience as a Data Engineer in production projects
Proficiency in Python (PySpark) and SQL for handling large datasets
Practical experience with cloud data processing platforms such as Microsoft Fabric, Azure Synapse, or Databricks
Familiarity with layered data architecture (medallion or similar) and Delta/Parquet formats
Experience working with Git in a team environment
Ability to integrate diverse sources, including Excel files from business users
Willingness to sign an NDA and undergo security verification
Fluent Polish
Nice-to-have requirements
DP-600 or DP-700 certification
Experience in the finance sector or other highly regulated environments
Knowledge of R (the model uses Python and R)
Previous work with CI/CD for data platforms
Familiarity with data quality frameworks
Experience with Agile and Scrum methodologies

