{"id":1162754,"url":"https://alion.io/job/factfinder-senior-site-reliability-engineer-sre-kubernetes-hybrid-cloud-mfd","title":"Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)","company":{"id":674168,"name":"FactFinder","domain":"fact-finder.com","url":"https://alion.io/company/fact-finder","size_band":"51-200","is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Join","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Berlin, Germany"],"countries":["DE"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":80000,"max_usd":166000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":25},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"ArgoCD","optional":false},{"name":"Kubernetes","optional":false},{"name":"KubeVirt","optional":false},{"name":"GitOps","optional":true},{"name":"Grafana","optional":true},{"name":"Incident Management","optional":true},{"name":"K3s","optional":true},{"name":"KEDA","optional":true},{"name":"kubeadm","optional":true},{"name":"OpenStack","optional":true},{"name":"Prometheus","optional":true},{"name":"Self-Healing","optional":true},{"name":"VLAN","optional":true}],"status":"live","first_seen_at":"2026-09-04T05:32:25Z","employer_posted_date":"2026-09-04","last_verified_at":"2026-09-24T00:45:32Z","board_verified":true,"closed_at":null,"days_open":20,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":20},"description":"Introduction\nAt a glance\nLocation & work model: Berlin, hybrid\nTech stack: Kubernetes on our own servers, Harvester (KubeVirt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph\nTeam: A growing SRE team - you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware\nProcess: Intro call · take-home task (~2h) · 90-min tech interview with our developers · leadership conversation · meet the team\nLanguages: Fluent English required; German is a plus, not a must\nWhy this role is special\nMost SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester - on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams - and you won't do it alone, a Team Lead SRE hire is coming next.\nSRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct - our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.\nYour first 90 days\nYou get to know both products, join the on-call rotation with a buddy, and own your first reliability topic - SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.\nYour mission\nDefine and own SLOs, SLIs and error budgets; drive data-informed reliability decisions\nLead incident response end-to-end: fast detection, clear communication, blameless postmortems - and reduce whole classes of incidents structurally, not case by case\nEliminate toil through automation and GitOps; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks\nHelp build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable - and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard\nPlan capacity, performance and cost across on-premises and cloud - including the large-catalogue and peak-season loads our merchants care about - and use AI tools wherever they measurably speed up diagnosis and operations\nYour profile\nMust-haves:\nKubernetes in production - built, not just used: you've set up and maintained clusters on your own servers (e.g. kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades - managed-only experience isn't enough for this role\nLived SRE practice: SLOs, error budgets, incident management, on-call\nHands-on experience with GitOpsor comparable infrastructure/deployment automation - experience with Argo CD or Flux is a strong plus\nSolid observability skills - metrics, logs, traces, alerting that people trust\nA strong automation instinct - you'd rather fix a problem's cause than repeat its workaround\nA collaborative, enabling mindset - you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution\nNice-to-haves (genuinely optional - we'll teach you the rest):\nHarvester, KubeVirt, vSphere/ESXi, OpenStack or similar virtualization/HCI platforms\nContainer storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)\nAuto-scaling (HPA, VPA, KEDA, cluster auto scaler) and capacity/cost planning\nExperience building Kubernetes operators/CRDs\nGerman language skillsCertifications (CKA, CKS) are welcome but no substitute for hands-on experience - in the tech interview we'll ask about what you've actually built and operated.\n\nYou don't tick every box - or your title was never “SRE”? Apply anyway. If you've owned production systems, handled incidents and worked deeply with Kubernetes, we want to hear from you - production experience and engineering mindset matter more to us than titles or buzzwords.\nTHE JOY OF WORKING WITH US\nImpact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.\nModern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right.\nAI-first mindset: We use AI as a real part of our daily work, not as a buzzword.\nOwnership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.\nFlexible work: Hybrid work model three office days per week with a focus on outcomes.\nStrong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.\nJob Location\nBerlin, Munich, Pforzheim or Stockholm (all Hybrid)","description_format":"text","description_chars":5047,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"German","level":"All levels","optional":false},{"language":"English","level":"Advanced (C1)","optional":false}]},"benefits":["Flexible schedule","Hybrid work"],"hiring_locations":[{"name":"Germany","iso":"DE","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","AI Search"],"lifecycle":[{"event":"open","at":"2026-09-24T00:45:32Z"}],"liveness":{"score":68,"band":"ok","label":"Likely open","p_open":1,"p_active":0.76,"p_room":0.9,"age_days":20,"expected_fill_days":42,"reasons":["conf:4","win:mid"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/factfinder-senior-site-reliability-engineer-sre-kubernetes-hybrid-cloud-mfd","json_url":"https://alion.io/job/factfinder-senior-site-reliability-engineer-sre-kubernetes-hybrid-cloud-mfd.json","meta":{"generated_at":"2026-09-25T01:36:23Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1644,"day_limit":5000,"remaining_today":3356,"minute_limit":60,"resets_at":"2026-09-26T00:00:00Z"}}}