{"id":1238328,"url":"https://alion.io/job/range-senior-infrastructure-operations-engineer-kubernetes-platform-reliability","title":"Senior Infrastructure & Operations Engineer (Kubernetes / Platform Reliability)","company":{"id":2047822,"name":"Range","domain":"range.org","url":"https://alion.io/company/range-org","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Deel","truth_index":{"grade":"B","score":75,"open_postings":3,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-27T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"inferred_company_offices","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":["London, United Kingdom"],"countries":["GB"],"hiring_countries":["CH"],"hiring_countries_total":1,"salary":{"min":75000,"max":130000,"currency":"USD","period":"year","gross":null,"usd_annual":130000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"CI/CD","optional":false},{"name":"Cloudflare","optional":false},{"name":"DNS","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"TCP/IP","optional":false},{"name":"ElasticSearch","optional":true}],"status":"live","first_seen_at":"2026-04-27T18:20:27Z","employer_posted_date":"2026-04-27","last_verified_at":"2026-09-25T23:57:51Z","board_verified":false,"closed_at":null,"days_open":153,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":153},"description":"We’re looking for a senior infrastructure and operations engineer to own and evolve our platform reliability. You’ll design, operate, and maintain our Kubernetes-based infrastructure, build reliable monitoring and alerting pipelines, and ensure our systems remain stable under real-world load and failure conditions. This is a hands-on role for someone with deep experience running production systems at scale and who focuses on making infrastructure predictable and stable. You’ll work across Kubernetes, networking, CI/CD, Cloudflare, and observability to create a platform engineers can trust.\nWhat You’ll Do\nDesign, deploy, and maintain production Kubernetes clusters.\n\nOwn cluster reliability, upgrades, security, and performance.\n\nBuild and operate monitoring, logging, and alerting pipelines.\n\nEnsure full-stack observability across infrastructure and services.\n\nDesign and maintain CI/CD pipelines that are fast, reproducible, and safe.\n\nImprove deployment strategies (rollouts, canaries, rollbacks).\n\nAutomate infrastructure provisioning and configuration.\n\nInvestigate and resolve production incidents.\n\nImprove system resilience, redundancy, and recovery strategies.\n\nDefine SLOs/SLIs and track reliability targets.\n\nOptimize and maintain our Cloudflare setup (caching, routing, security, edge behavior).\n\nWork closely with engineering teams to improve operational practices.\n\nIdentify and remove single points of failure.\n\nWhat We’re Looking For\nMust-have\nSenior-level experience operating production infrastructure.\n\nDeep, hands-on expertise with Kubernetes (cluster internals, networking, storage, security).\n\nStrong networking fundamentals (TCP/IP, routing, DNS, TLS, load balancing).\n\nExperience debugging distributed systems and network-related issues.\n\nExperience optimizing CDN and edge setups, including Cloudflare.\n\nStrong experience building monitoring and observability systems.\n\nExperience with metrics, logs, traces, and alerting pipelines.\n\nExperience designing reliable CI/CD pipelines.\n\nStrong Linux fundamentals.\n\nExperience with infrastructure as code and automation.\n\nAbility to debug issues across the entire stack.\n\nExperience handling incidents and conducting postmortems.\n\nNice-to-have\nExperience with multi-cluster or multi-region setups.\n\nExperience with high-throughput or data-heavy systems.\n\nExperience with Elasticsearch or large-scale data infrastructure.\n\nExperience with service meshes.\n\nExperience with cost optimization and capacity planning.\n\nExperience in regulated or reliability-focused environments.\n\nHow You Work\nYou assume infrastructure will fail and design accordingly.\n\nYou prioritize reliability, visibility, and recoverability.\n\nYou build systems that engineers trust in production.\n\nYou automate carefully and deliberately.\n\nYou are calm and methodical during incidents.\n\nYou focus on long-term stability over short-term fixes.\n\nYou document and standardize important processes.\n\nExample Problems You Might Work On\nHardening Kubernetes clusters for high availability and safe upgrades.\n\nDebugging network latency or connectivity issues across services.\n\nOptimizing Cloudflare caching, routing, and edge security rules.\n\nBuilding monitoring pipelines that provide reliable signals.\n\nDesigning alerting that is actionable and low-noise.\n\nImproving deployment reliability and rollback safety.\n\nRemoving single points of failure in production systems.\n\nEnsuring observability across all services and data pipelines.\n\nWhy Join Range\nJoin one of the fastest-growing sectors in Web3 as stablecoins reach mass adoption.\n\nCompetitive compensation with meaningful equity upside.\n\nStrong potential for growth and leadership opportunities.\n\nRemote-first culture with bi-yearly international off-sites.\n\nOpportunities for global conference travel and ecosystem engagement.\n\nHealth and wellness benefits.\n\nHow to Apply\nSend us:\nA short introduction and your background.\n\nExamples of infrastructure or platform work you’ve led.\n\nAny public write-ups, repositories, or talks (if available).\n\nWe’re particularly interested in engineers who have built and operated Kubernetes platforms, improved network reliability, and optimized CDN/edge setups such as Cloudflare in production.","description_format":"text","description_chars":4221,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity"],"hiring_locations":[{"name":"Switzerland","iso":"CH","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Blockchain & Crypto","Cryptocurrencies","Stablecoins"],"lifecycle":[{"event":"open","at":"2026-09-25T16:23:32Z"}],"liveness":{"score":5,"band":"cold","label":"Long shot","p_open":1,"p_active":0.177,"p_room":0.28,"age_days":152,"expected_fill_days":30,"reasons":["conf:29","win:tail","crowd:"],"computed_at":"2026-09-27T05:45:00Z"},"pay":{"stated_usd_annual":130000,"is_top_pay":false},"html_url":"https://alion.io/job/range-senior-infrastructure-operations-engineer-kubernetes-platform-reliability","json_url":"https://alion.io/job/range-senior-infrastructure-operations-engineer-kubernetes-platform-reliability.json","meta":{"generated_at":"2026-09-28T04:56:18Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3216,"day_limit":5000,"remaining_today":1784,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}