615,181open jobs
29,638companies
85,681added this week
Browse all
Salary
$69k – $103k per year (Estimated)
Location
Remote (Poland)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Sigma Software is a technology services company founded in 2002 as the Ukrainian development arm of the Swedish Sigma Group, and now a global organisation with delivery centres in Ukraine, Poland, Sweden, the United States and Latin America. It builds custom software for clients in advertising technology, automotive, aviation, gaming, telecommunications and the public sector, and runs its own product and startup arms alongside the services business. The company is known for keeping its Ukrainian engineering operation running through the war and for a substantial programme of technology work supporting Ukrainian defence and public institutions.

Join Sigma Software to build large-scale data infrastructure powering a real-time AdTech platform processing hundreds of millions of auction requests daily. We are looking for a Senior Data Engineer who enjoys solving complex distributed data challenges and building production-grade ML-oriented data systems.

You will become part of a dedicated Sigma Software team developing predictive modeling and optimization capabilities for a live advertising ecosystem. The role combines large-scale event processing, streaming and batch pipelines, experimentation infrastructure, and high-throughput data engineering in a cloud-native environment.

We as a company offer the opportunity to work on impactful global products, collaborate with experienced engineers, and contribute to architecture decisions while growing your expertise in large-scale distributed systems and modern data platforms.

CUSTOMER

Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a large-scale ad exchange handling hundreds of millions of auction requests per day and is actively investing in predictive decisioning technologies to optimize advertising outcomes in real time.

PROJECT

The project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform performs real-time supply scoring and filtering, contextual performance estimation, look-alike audience generation, and multi-objective optimization under business constraints.

The solution processes massive-scale event and auction datasets and includes feature engineering pipelines, streaming and batch ingestion, experimentation infrastructure, point-in-time-correct training data generation, and ML-oriented data services with strict operational reliability and compliance requirements.

  • Write and defend diagnostic SQL queries against large-scale production datasets
  • Build and maintain ingestion pipelines for bid, win, and impression logs into BigQuery
  • Harmonize fields across independently designed datasets and maintain versioned field mappings
  • Develop point-in-time-correct feature tables and aggregation pipelines
  • Design and maintain conversion and labeling pipelines with delayed label handling
  • Own the data serving write path, schema contracts, publishing flows, and freshness SLOs
  • Build experimentation infrastructure including traffic splitting and reporting pipelines
  • Perform large-scale historical backfills and safe reprocessing after mapping changes
  • Implement data isolation and safe-aggregation controls for advertiser data protection
  • Develop automated data quality validation frameworks
  • Collaborate closely with Customer engineers and prepare operational documentation
  • Contribute to architecture discussions and platform scalability improvements
  • 5+ years of experience in Data Engineering
  • At least 2 years of experience working with production ML or large-scale analytics pipelines
  • Expert-level SQL skills including window functions and incremental processing patterns
  • Strong Python skills for production-grade pipeline development
  • Hands-on experience with Spark or PySpark
  • Experience designing ETL / ELT pipelines with Airflow, Cloud Composer, Dagster, or similar tools
  • Experience working with cloud data warehouses at scale, preferably BigQuery
  • Strong understanding of data modeling and point-in-time correctness
  • Experience working with event-driven or clickstream datasets at very large scale
  • Experience supporting business-critical production pipelines
  • Upper-Intermediate English level or higher

WILL BE A PLUS

  • Experience with GCP services including Dataflow, Pub/Sub, GCS, and Beam
  • Experience building streaming or near-real-time ingestion systems
  • Understanding of feature stores, train/serve skew, and label leakage prevention
  • Experience in AdTech or auction-based environments
  • Experience handling delayed or incomplete labels in ML systems
  • Experience with dbt or similar transformation frameworks
  • Experience delivering solutions into Customer-owned infrastructure
  • Knowledge of GDPR/CCPA-related privacy engineering practices
  • Experience with experimentation infrastructure and statistical validation pipelines
  • Experience working in hybrid cloud/on-prem Linux environments
  • Terraform and Kubernetes experience
  • Experience optimizing warehouse cost and performance

PERSONAL PROFILE

  • Strong analytical and problem-solving skills
  • Ownership-oriented mindset
  • Ability to work independently in a client-facing environment
  • Strong communication and documentation skills
  • Comfortable working in a fast-paced engineering environment
  • Collaborative and proactive attitud
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
615,181 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Warsaw
$81k – $157k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Ireland
Python
JavaScript
PowerShell
DevOps
GCP
Azure
CI/CD
AWS
Platform Engineering
IAM
Cybersecurity
Okta
CyberArk
Least Privilege
Microsoft Entra ID
Apply
$158k – $231k per year • Equity • In office • Full-Time • 8+ years exp • Toronto • San Francisco • Denver
Python
SQL
Apply
Software Developer 1 day ago
$22k – $60k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Athens
Python
Java
SQL
C++
Python
Flask
Databases
MySQL
RabbitMQ
ElasticSearch
Apache Kafka
OpenSearch
DevOps
CI/CD
GitLab
Apply
$116k – $173k per year • In office • Full-Time • 4+ years exp • PhD • Cambridge
Python
R
R
ggplot2
Bioconductor
AI/ML
Copilot
Claude Code
Scikit-learn
Prompt Engineering
AI Agents
TensorFlow
PyTorch
LLM
RAG
Structured Outputs
Agentic Workflows
Multi-Agent Systems
DevOps
GCP
SLURM
Git
AWS
GitHub
HPC
Apply
$48k – $94k per year (Estimated) • In office • Part-Time • Tampa
Python
SQL
AI/ML
Prompt Engineering
NumPy
DevOps
Git
Apply
$65k – $115k per year (Estimated) • Remote • Full-Time • 8+ years exp • Kraków
Python
Java
SQL
Scala
Databases
Apache Iceberg
Trino
AI/ML
Spark
Flink
DevOps
CI/CD
AWS
Platform Engineering
Apply
$59k – $130k per year (Estimated) • Remote • Full-Time • 8+ years exp • Kyiv
Python
Java
SQL
Scala
Databases
Apache Iceberg
Trino
AI/ML
Spark
Flink
DevOps
CI/CD
AWS
Platform Engineering
Apply
$51k – $127k per year (Estimated) • Remote • Full-Time • 7+ years exp • São Paulo
Java
SQL
Java
Spring Boot
Databases
Apache Kafka
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Amazon EKS
Apply
$63k – $107k per year (Estimated) • Remote • Full-Time • 5+ years exp • PhD • Warsaw
Python
Databases
Redis
Aerospike
Google Bigtable
AI/ML
MLFlow
Vertex AI
Kubeflow
Feature Store
DevOps
Terraform
GCP
CI/CD
Docker
Kubernetes
Platform Engineering
Argo Workflows
Cybersecurity
ISO 27001
SOC 2
GDPR
Apply
$67k – $100k per year (Estimated) • Remote • Full-Time • Warsaw
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Apache Kafka
Amazon Redshift
AI/ML
Spark
DevOps
Terraform
Azure
Docker
Kubernetes
Analytics
ETL/ELT
Apply
$56k – $149k per year (Estimated) • Remote/Hybrid • Internship • Warsaw
Python
AI/ML
LangGraph
LangChain
Embeddings
Prompt Engineering
AI Agents
LLM
RAG
Semantic Search
Semantic Search
DevOps
Rest API
GCP
Azure
Git
AWS
GitHub
Apply
Account Specialist 1 day ago
Remote/Hybrid • Full-Time • Warsaw
Apply
$47k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Warsaw
Apply
$61k – $112k per year (Estimated) • In office • Full-Time • Master's Degree • Warsaw
Apply
$39k – $92k per year (Estimated) • Remote/Hybrid • Full-Time • Warsaw
Marketing
Salesforce
Apply
See all jobs
This is one of many
615,181 more open roles from verified company boards, updated every day.