723,472open jobs
43,214companies
103,113added this week
Browse all
Salary
$79k – $163k per year (Estimated)
Location
Remote/Hybrid (Berlin, Germany)
Seniority
Senior
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 4, 2026.

Overview
Company
Impact
Profile match
Help every visitor find the right product, fast. AI-powered search and product discovery that turns eCommerce traffic into sales and loyalty.

Introduction

FACT-Finder entwickelt Product-Discovery-Technologie für den eCommerce und ist mit den Produkten Next Generation und Infinity bei führenden Online-Shops in Europa im Einsatz. Beide Produkte bewegen sich aktuell in Richtung einer modernen, hybriden Plattform auf Basis von Kubernetes und Harvester - mit der Option, mittelfristig vollständig in die Cloud zu skalieren. Als Senior Site Reliability Engineer (SRE) sorgst du dafür, dass unsere Systeme über diese Transformation hinweg schnell, verfügbar und skalierbar bleiben. Du arbeitest mit dem Hosting-Team und erfahrenen Engineers zusammen und gestaltest den Weg zu einer modernen SaaS-Company aktiv mit.

Deine Aufgaben

  • Du definierst und verantwortest SLOs, SLIs und Error Budgets über beide Produkte hinweg und triffst datenbasierte Entscheidungen zu Reliability und Performance.
  • Du treibst Incident Response voran: schnelle Detektion, klare Kommunikation, blameless Postmortems und nachhaltige Follow-ups.
  • Du reduzierst manuelle Arbeit konsequent durch Automatisierung und GitOps (z. B. Argo CD / Flux) und baust self-healing sowie self-service Fähigkeiten aus.
  • Du unterstützt den Aufbau eines NG Search Operators (Custom Kubernetes Operator / CRDs) und die Einführung von Auto-Scaling (HPA, VPA, KEDA, Cluster Autoscaler).
  • Du entwickelst unsere Observability weiter - Metrics, Logs, Traces, Alerting und Runbooks, die on-call wirklich helfen.
  • Du planst Kapazität und Kosten über On-Premise (Frankfurt, Stockholm) und Cloud hinweg - inklusive Burst-Szenarien in die Public Cloud.
  • Du nutzt AI-Tools, um Diagnose, Alerting und operative Workflows spürbar zu

Dein Profil

  • Erfahrung als SRE, Infrastructure oder Production Engineer in einem SaaS- oder Plattform-Umfeld - oder ein starker Software-/Operations-Hintergrund mit klarem Willen, in die SRE-Rolle hineinzuwachsen.
  • Solides Verständnis von SLOs, Error Budgets, Incident Management und Observability.
  • Hands-on Erfahrung mit Kubernetes und Interesse an Cluster-Lifecycle, Upgrades und Operator-Pattern.
  • Erfahrung oder starkes Interesse an Harvester bzw. vergleichbaren HCI-/Virtualisierungsplattformen (KubeVirt, vSphere/ESXi, OpenStack).
  • Vertrautheit mit GitOps (Argo CD / Flux), Container Storage (Longhorn, Ceph) und Kubernetes Networking (Load Balancing, Ingress).
  • Kenntnisse zu Auto-Scaling-Primitiven (HPA, VPA, Cluster Autoscaler, KEDA) und Kapazitätsplanung on-prem und in der Cloud.
  • Verständnis für Netzwerke in produktionsnahen Rechenzentren (u. a. VLAN).
  • Ausgeprägter Automatisierungsinstinkt und eine Haltung, Toil strukturell zu eliminieren.
  • Praktische Erfahrung im Einsatz von AI-Tools im operativen Betrieb.
  • Sehr gute Englischkenntnisse; Deutsch von Vorteil.

THE JOY OF WORKING WITH US

  • Impact from day one: Deine Arbeit wirkt direkt auf die Umsätze führender eCommerce-Marken in Europa.
  • Moderner Tech-Stack: Kubernetes, Harvester, GitOps, Auto-Scaling und ein spannender Weg in Richtung Cloud - mit Raum, Dinge neu und richtig zu bauen.
  • AI-first Mindset: Wir nutzen AI nicht als Buzzword, sondern als festen Bestandteil unserer täglichen Arbeit.
  • Ownership & Wachstum: Klare Verantwortung, kurze Entscheidungswege und die Möglichkeit, deine Rolle aktiv mitzugestalten.
  • Flexibles Arbeiten: Hybrides Arbeitsmodell mit Fokus auf Ergebnisse.
  • Starkes Team: Erfahrene Engineers, offene Feedback-Kultur und ein Umfeld, in dem Reliability als Engineering-Disziplin ernst genommen wird.

Standort

Berlin, München, Pforzheim oder Stockholm (Hybrid)

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
723,472 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Berlin
$93k – $173k per year (Estimated) • Remote • Full-Time • London
Python
JavaScript
Ruby
Node JS
Ruby
Ruby on Rails
Devise
Databases
PostgreSQL
DevOps
Terraform
GCP
OpenTofu
Terragrunt
Pulumi
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
Platform Engineering
Analytics
ETL/ELT
Management
Agile
Apply
Platform Engineer 27 min ago
$77k per year • Remote • Full-Time • London
Python
JavaScript
Ruby
Node JS
Databases
PostgreSQL
DevOps
Terraform
GCP
OpenTofu
Terragrunt
Pulumi
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
Platform Engineering
Analytics
ETL/ELT
Management
Agile
Apply
$21k – $45k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
JavaScript
Java
SQL
Frontend
React.js
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Linux
Unix
Management
Agile
Apply
$32k – $76k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
Python
Java
SQL
Databases
Databricks
AI/ML
Claude
Spark
Machine Learning
DevOps
CI/CD
AWS
Docker
Kubernetes
Analytics
Power BI
Looker
Apply
Consultant-1 11 min ago
$18k – $51k per year (Estimated) • In office • Full-Time • 4+ years exp • Bengaluru
Python
SQL
Python
Flask
FastAPI
Databases
Databricks
AI/ML
LangGraph
LangChain
Spark
XGBoost
Fine-tuning
Embeddings
Scikit-learn
NLP
Langflow
LightGBM
TensorFlow
Pandas
PyTorch
LLM
RAG
Flowise
LLMOps
Human-in-the-Loop
DevOps
Azure
CI/CD
Docker
Kubernetes
Incident Management
Apply
$80k – $166k per year (Estimated) • Remote/Hybrid • Full-Time • Berlin
DevOps
K3s
KEDA
Prometheus
GitOps
ArgoCD
Kubernetes
Grafana
Self-Healing
kubeadm
OpenStack
KubeVirt
Incident Management
VLAN
Apply
$64k – $123k per year (Estimated) • Remote/Hybrid • Full-Time • Berlin
Management
Slack
Notion
Service Desk
Marketing
Salesforce
LinkedIn
Apply
$97k – $197k per year (Estimated) • Remote/Hybrid • Full-Time • Berlin
DevOps
KEDA
GitOps
ArgoCD
Kubernetes
Platform Engineering
OpenStack
KubeVirt
Incident Management
VLAN
Apply
$42k – $103k per year (Estimated) • In office • Full-Time • Berlin
Management
Slack
Notion
Marketing
Salesforce
Apply
Product Manager 12 hours ago
$66k – $144k per year (Estimated) • In office • Full-Time • Berlin • Pune • Lisbon • London
Management
Agile
Apply
$61k – $112k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Berlin
Python
JavaScript
TypeScript
Python
FastAPI
AI/ML
Cursor
Claude Code
LLM
Frontend
Svelte
DevOps
Rest API
Management
Agile
Apply
$57k – $156k per year (Estimated) • Remote • Part-Time • Berlin
Apply
$80k – $182k per year (Estimated) • In office • Berlin
Python
SQL
Databases
Snowflake
AI/ML
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Amazon EKS
Google GKE
Azure AKS
Apply
$62k – $131k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Berlin
Apply
See all jobs
This is one of many
723,472 more open roles from verified company boards, updated every day.