We are hiring on behalf of an established international technology company whose core platform processes high-throughput, real-time data at massive global scale.
As aSenior Data Platform Engineer, you will take operational ownership of a self-hosted, open-source data infrastructure running on Kubernetes and private cloud systems. Working at the intersection of Platform Engineering and SRE, your focus will be on platform reliability, automation, GitOps practices, and supporting high-performance data engineering and analytics teams.
This is a hands-on role at the intersection of Platform Engineering and Site Reliability Engineering. You will work closely with architects and technical leaders, contribute to architectural decisions, and act as a key point of contact for the platform’s reliability and evolution. This is not a traditional data engineering, analytics, or ticket-based operations role.
Key takeaways:
Stack:
Orchestration & Infrastructure: Kubernetes, Terraform, Helm, GitOps (ArgoCD/Flux), Private Cloud (OpenStack / AWS).
Data Systems: Event Streaming (Kafka/Redpanda), Analytical Storage (Apache Druid/Elasticsearch), Data Processing (Spark, Airflow), PostgreSQL.
Salary: 23 000 -26 600 PLN gross monthly on an employment contract (open to negotiations)
Working model: 2 days per week in the Krakow office; remote work within Poland may be considered for an excellent match
Location: Krakow, Podgórze,
Recruitment process:
- A call with Motife recruiter
- Short online coding assessment followed by two technical and system-design interviews
- Final meeting with the CTO
Responsibilities:
Platform Ownership
- Own the operational health and ongoing development of a self-hosted data platform running on Kubernetes and OpenStack.
- Operate production systems including Apache Druid, Redpanda, Spark, Airflow, PostgreSQL, and Elasticsearch.
- Participate in platform stabilisation, upgrades, capacity planning, and ongoing performance improvements as the platform scales.
- Build reliable operating patterns for stateful Kubernetes workloads, including storage, replication, disruption management, and rolling upgrades.
Reliability and Incident Management
- Define and monitor the KPIs, alerts, dashboards, SLOs, and error budgets required to maintain platform reliability.
- Build observability coverage using Prometheus, Grafana, logging, and distributed tracing.
- Participate in the on-call rotation, respond to production incidents, and coordinate the resolution of major platform issues.
- Drive root-cause analysis, post-incident reviews, runbooks, and preventive improvements.
Infrastructure and Automation
- Codify infrastructure and platform configuration using Terraform, Helm or Kustomize, and GitOps practices.
- Improve CI/CD processes to make platform changes safe, repeatable, and easy for the wider team to support.
- Automate recurring operational work and reduce manual intervention across the platform.
- Apply appropriate security, secrets management, backup, and disaster-recovery practices.
Architecture and Collaboration
- Contribute to architectural discussions and help shape the future technology stack rather than simply maintaining the existing design.
- Partner with globally distributed architects, SREs, data engineers, analysts, and data scientists.
- Translate platform requirements into scalable and supportable technical solutions.
- Use AI-assisted engineering tools pragmatically while maintaining strong quality controls for production infrastructure.
Requirements:
Platform and SRE Expertise
- Strong hands-on experience in Platform Engineering, Site Reliability Engineering, DevOps, or infrastructure engineering.
- Advanced knowledge of Kubernetes in production, particularly stateful workloads, persistent storage, disruption management, and safe upgrades.
- Production experience with Infrastructure as Code and configuration management using tools such as Terraform, Helm, or Kustomize.
- Practical experience with GitOps and automated delivery, ideally using ArgoCD, Flux, or comparable tooling.
- Experience owning production systems, participating in on-call rotations, and resolving complex incidents.
Data Platform Experience
- Hands-on experience operating Kafka, Redpanda, or another production event-streaming platform.
- Experience administering or supporting at least one additional data technology such as Spark, Airflow, PostgreSQL, Elasticsearch, Hadoop, or Apache Druid.
- Good understanding of data ingestion, streaming, partitioning, replication, consumer lag, capacity, and performance.
- Experience implementing monitoring, alerting, dashboards, and structured root-cause analysis.
- You do not need experience with every technology in the stack. Apache Druid, Spark, Redpanda, and OpenStack can be learned if you bring strong Kubernetes, platform reliability, and data infrastructure fundamentals.
Collaboration and Mindset
- Strong ownership mindset and the ability to operate independently in a small, high-impact team.
- Confidence contributing to architectural discussions and challenging existing solutions when improvements are possible.
- Openness to learning unfamiliar technologies and using modern AI-assisted development tools.
- Clear written and spoken English
Apply now
Join a high-impact engineering team and take ownership of the reliability and future direction of a large-scale, self-hosted data platform. If you are strongest in Kubernetes, Platform Engineering, or SRE and want to deepen your expertise in data infrastructure, we would like to hear from you.

