First seen by Alion on Oct 5, 2026.
We are looking for a skilled Data Engineer to join an international technology and data team on a 6-month engagement, starting in October.
This is a hands-on role focused on building data solutions for the insurance domain. You will use Databricks AI capabilities and Apache Spark to extract, process and structure information from both typed and handwritten insurance documents. A key part of the role will be transforming and modelling this data so it can be reliably used for machine learning and analytics use cases. You will be working 100% remotely in international team.
Key responsibilities
Design and develop scalable data pipelines using Azure, Databricks, Python and PySpark.
Use Databricks AI tooling and Spark-based processing to extract and process data from typed and handwritten insurance documentation.
Clean, validate, transform and model data for machine learning, analytics and BI use cases.
Build reliable datasets that can support downstream reporting, analytical models and AI/ML solutions.
Develop and maintain data transformations using Python and PySpark.
Support end-to-end data engineering and analytics initiatives, from data ingestion through processing, quality validation and delivery.
Apply clean-code principles, write maintainable and well-documented code, and contribute to code reviews.
Create and maintain unit tests to ensure the quality, reliability and stability of data pipelines.
Collaborate with data, analytics and business stakeholders to understand requirements and translate them into practical technical solutions.
Identify data-quality issues and contribute to improvements in data governance, consistency and usability.
Ideal candidate profile
Commercial experience as a Data Engineer or in a similar data-focused technical role.
Strong hands-on experience with Microsoft Azure and Databricks.
Very good knowledge of Python and practical experience with PySpark.
Experience building and maintaining data pipelines using Apache Spark and Databricks.
Experience extracting, processing, cleansing and transforming structured and unstructured or semi-structured data.
Understanding of data modelling principles and the ability to prepare data for machine learning, advanced analytics and BI/reporting use cases.
Experience writing unit tests and validating data pipeline logic.
Strong focus on clean code, code quality, maintainability and technical documentation.
Fluent English at least B2
Availability to start in October (ASAP)
Availability for a 6-month contract.
Conditions
Form of cooperation: B2B contract
Salary: 160 - 220 PLN net/h
100% remote work – you can work from any place in Poland
Start: ASAP - less than 2 weeks notice period
Project for 6 months - possible to be extended
Recruitment steps
Online meeting or phone call with Recruiter from KUBO (20-30 min.)
Technical verification call focused on your experience and project requirements
Technical interview with the client, including practical questions and potentially live coding or a technical task
Feedback and decision
Tech stack
- Data Engineering
- Azure
- Databricks
- Python
- PySpark
- Apache Spark
- Data modeling
- unit tests
- Machine Learning
- data governance

