{"id":1061753,"url":"https://alion.io/job/factfinder-team-lead-site-reliability-engineering-all-genders","title":"Team Lead - Site Reliability Engineering (all genders)","company":{"id":674168,"name":"FactFinder","domain":"fact-finder.com","url":"https://alion.io/company/fact-finder","size_band":"51-200","is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Join","truth_index":null},"role":"Leadership","role_family":"Leadership","seniority":"lead","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Berlin, Germany"],"countries":["DE"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":97000,"max_usd":197000,"period":"year","method":"role_seniority_country_cell","sample_n":8},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"ArgoCD","optional":false},{"name":"GitOps","optional":false},{"name":"Incident Management","optional":false},{"name":"KEDA","optional":false},{"name":"Kubernetes","optional":false},{"name":"KubeVirt","optional":false},{"name":"OpenStack","optional":false},{"name":"Platform Engineering","optional":false},{"name":"VLAN","optional":false}],"status":"live","first_seen_at":"2026-09-01T07:33:46Z","employer_posted_date":"2026-09-01","last_verified_at":"2026-09-24T00:45:32Z","board_verified":true,"closed_at":null,"days_open":23,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":23},"description":"Introduction\nFACT-Finder builds product discovery technology for eCommerce and is trusted by leading online shops across Europe with its two products Next Generation and Infinity. We are actively modernizing our hosting toward Kubernetes on Harvester - as an on-prem hybrid with the option to scale fully into the cloud in the mid-term. As Team Lead Site Reliability Engineering (all genders), you own the reliability, scalability, and cost of our hosting environments, drive this transformation end-to-end, and lead the team that delivers it.\nYour mission\nYou own the operational health of our hosting across on-premise (Frankfurt, Stockholm) and cloud - availability, performance, and incident management.\nYou actively drive the modernization toward Kubernetes on Harvester: cluster topology, storage (Longhorn), networking (VLAN, load balancing, ingress), backup, and disaster recovery.\nYou build a production-grade k8s platform: lifecycle, upgrades, RBAC, secrets, GitOps (Argo CD / Flux), observability, and policy guardrails.\nYou shape the NG Search Operator (custom Kubernetes operator) and solve auto-scaling (HPA, VPA, KEDA, cluster autoscaler) for the current architecture.\nYou concretely define our on-prem hybrid model: which workloads run where, how we burst into the cloud, how we keep latency and cost under control - while keeping the architecture portable enough for a future cloud-only move.\nYou own capacity planning and hosting cost and turn cost into a deliberate, managed lever.\nYou lead and develop our currently 4-person Hosting team, own performance and technical direction, and set the standards and ownership culture.\nYou make AI a core part of our operations: diagnosis, automation, monitoring, and insight.\nYour profile\nStrong background in infrastructure or platform engineering across on-premise and cloud.\nHands-on depth with Kubernetes in production: cluster lifecycle, upgrades, networking, storage, RBAC, observability, GitOps delivery.\nProven people leadership experience, excellent communication and stakeholder management skills.\nIdeally practical experience with Harvester or comparable HCI/virtualization platforms (KubeVirt, vSphere/ESXi, OpenStack).\nExperience leading a real migration from bare metal / classic VMs to a k8s-based platform - including stateful workloads, storage migration, cutover, and rollback.\nComfort designing or operating Kubernetes operators (custom controllers / CRDs), ideally for stateful systems like search, databases, or streaming.\nSolid grasp of auto-scaling primitives (HPA, VPA, cluster autoscaler, KEDA) and how they interact with capacity planning on-prem and in the cloud.\nExperience with on-prem hybrid architectures and owning reliability, capacity, and cost for production systems.\nHands-on fluency with AI tools in day-to-day operations.\nFluent English; German is a plus.\nTHE JOY OF WORKING WITH US\nImpact from day one: Your work directly influences the revenue of leading eCommerce brands across Europe.\nLeadership with real scope: You lead an established team and shape our platform in a decisive phase of our transformation.\nModern tech stack: Kubernetes, Harvester, GitOps, auto-scaling, and an exciting path toward the cloud - with room to build things right.\nAI-first mindset: We use AI as a real part of our daily work, not as a buzzword.\nOwnership & growth: Clear responsibility, short decision paths, and the opportunity to actively shape your role.\nFlexible work: Hybrid work model with a focus on outcomes.\nStrong team: Experienced engineers, an open feedback culture, and an environment where reliability is treated as a real engineering discipline.\nAttractive benefits: Competitive salary, modern equipment, learning budget, and regular team events.\nJob Location\nBerlin, Munich, Pforzheim or Stockholm (hybrid)","description_format":"text","description_chars":3805,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"Advanced (C1)","optional":false}]},"benefits":["Flexible schedule","Hybrid work"],"hiring_locations":[{"name":"Germany","iso":"DE","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","AI Search"],"lifecycle":[{"event":"open","at":"2026-09-19T09:00:49Z"}],"liveness":{"score":61,"band":"ok","label":"Likely open","p_open":1,"p_active":0.676,"p_room":0.9,"age_days":22,"expected_fill_days":35,"reasons":["conf:4","win:mid"],"computed_at":"2026-09-24T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/factfinder-team-lead-site-reliability-engineering-all-genders","json_url":"https://alion.io/job/factfinder-team-lead-site-reliability-engineering-all-genders.json","meta":{"generated_at":"2026-09-24T16:31:53Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}