{"id":2036061,"url":"https://alion.io/job/nexcess-platform-sre","title":"Platform SRE","company":{"id":2155317,"name":"Nexcess","domain":"nexcess.com","url":"https://alion.io/company/nexcess-com","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Rippling","truth_index":{"grade":"B","score":75,"open_postings":4,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-10T05:45:15Z"}},"role":"DevOps","role_family":"DevOps","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":11500,"max_usd":30000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":16},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"Bash","optional":false},{"name":"Chef","optional":false},{"name":"CI/CD","optional":false},{"name":"Configuration Management","optional":false},{"name":"Datadog","optional":false},{"name":"Docker","optional":false},{"name":"Grafana","optional":false},{"name":"Human-in-the-Loop","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Puppet","optional":false},{"name":"Python","optional":false},{"name":"Self-Healing","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"SRE","optional":false},{"name":"Terraform","optional":false},{"name":"Progressive Delivery","optional":true}],"status":"live","first_seen_at":"2026-10-07T15:54:24Z","employer_posted_date":"2026-10-07","last_verified_at":"2026-10-10T23:43:47Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"About Nexcess\nNexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.\nThe Platform SRE is a member of the Platform SRE team, the group responsible for the operational health, resiliency, and modernization of our hosting and cloud infrastructure. This role sits at the intersection of systems engineering, site reliability engineering, and automation : combining hands-on technical debt remediation with large-scale automation and reliability practice.\nThe Platform SRE team owns three core mandates:\n1. Technical debt reduction : systematically identifying and remediating aging firmware, kernels, operating systems, and software across the managed hosting fleet and managed application environments.\n2. Automation of critical operational workflows : provisioning, patching, remediation, and release processes across both managed apps and managed hosting fleets, with automated release planning and execution under human supervision (not “automation for automation’s sake,” but automation with a human checkpoint before production impact).\n3. Reliability and incident support : defining, instrumenting, and tracking SLIs/SLOs for platform engineering and operations, visualized through dashboards and reporting tools, and providing fast, expert frontline response and remediation during service-impacting events (incident command and process ownership sit with the Incident Management team; this team is the technical responder, not the incident owner).\nThis position serves as a senior technical point of contact for platform engineering, driving initiatives around scalability, fault tolerance, automation, and operational excellence across production infrastructure.\nKey Responsibilities\nTechnical Debt & Platform Modernization\nOwn the lifecycle of firmware, kernel, OS, and software patching across the managed hosting fleet and managed application environments\nBuild a standing inventory and risk model of technical debt (end-of-life OS versions, unpatched firmware, deprecated software) and drive prioritized remediation plans\nEvaluate and implement infrastructure modernization initiatives, replacing manual or legacy processes with supportable, automated alternatives\nAutomation & Release Engineering\nDesign and build automation for provisioning, deployment, patching, remediation, and configuration management across managed apps and managed hosting fleets\nOwn the design of automated release pipelines : planning, staging, and executing releases with defined human-in-the-loop approval gates\nDevelop self-healing and auto-remediation capability for common failure modes to reduce manual operational load\nSupport and extend CI/CD workflows and infrastructure-as-code practices across the platform\nReliability Engineering, SLIs/SLOs & Observability\nDefine SLIs and SLOs for platform engineering and operations in partnership with engineering and product stakeholders\nInstrument systems to measure SLIs accurately and build SLO tracking into standard reporting\nBuild and maintain dashboards (e.g., Grafana, Datadog, or equivalent visualization tooling) to make SLI/SLO performance, error budgets, and platform health visible to engineering and leadership\nContinuously improve platform observability : monitoring, alerting, logging, and tracing : across distributed and containerized environments\nIncident Response & Remediation (Support Role)\nServe as the frontline technical responder: acknowledge pages quickly, diagnose, and remediate platform-level issues\nPartner with the Incident Management team, who own incident command, severity classification, and customer communication : this role provides the technical hands and expertise, not incident ownership\nContribute technical findings to blameless root cause analysis (RCA) and own follow-through on corrective actions for platform systems\nMaintain runbooks and on-call readiness for platform and infrastructure systems\nTrack incident trends on platform systems and feed them back into the technical debt and automation roadmap\nCollaboration & Technical Leadership\nPartner with software engineering teams on platform architecture, operational readiness reviews, and scalability initiatives\nSupport platform security, compliance, and operational governance requirements\nMentor engineers and contribute to technical leadership and knowledge-sharing across the team\nMaintain clear operational documentation and contribute to team standards and process improvement\nOther duties as assigned\nRequirements\n3-5+ years of experience in platform engineering, systems engineering, SRE, or infrastructure operations (level based on experience and scope)\nAdvanced Linux systems administration and troubleshooting expertise, including kernel and firmware-level familiarity\nStrong experience with Kubernetes, Docker, and container orchestration/distributed systems\nHands-on automation and infrastructure-as-code experience (e.g., Terraform, Ansible, Puppet/Chef, or equivalent)\nExperience building or maintaining CI/CD and automated release/deployment pipelines\nExperience defining and tracking SLIs/SLOs and working with observability/visualization tools (e.g., Grafana, Datadog, Prometheus, or equivalent)\nExperience supporting enterprise-scale, high-concurrency, or customer-impacting production environments\nDemonstrated experience as a technical responder in production incidents, including root cause analysis and corrective action follow-through\nStrong scripting ability (e.g., Python, Bash, Go) for automation and tooling\nStrong troubleshooting skills across compute, network, storage, and application layers\nExperience supporting cloud-hosted, managed hosting, or hybrid infrastructure environments\nAbility to lead technical initiatives, mentor others, and communicate clearly across teams\nPreferred Qualifications\nExperience owning fleet-wide firmware/OS patch management programs at scale\nExperience designing human-in-the-loop release automation or progressive delivery systems (canary, blue/green)\nFamiliarity with error budgets and SLO-driven prioritization frameworks\nExperience with configuration/patch management at scale across heterogeneous hardware fleets","description_format":"text","description_chars":6527,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-10-07T16:01:23Z"}],"visa":[],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":25,"reasons":["conf:2","velocity","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/nexcess-platform-sre","json_url":"https://alion.io/job/nexcess-platform-sre.json","meta":{"generated_at":"2026-10-11T01:55:06Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2890,"day_limit":5000,"remaining_today":2110,"minute_limit":60,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2155317},"rest":"https://alion.io/mcp/rest/get_company?id=2155317"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fnexcess-platform-sre"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fnexcess-platform-sre"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fnexcess-platform-sre"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/nexcess-platform-sre\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fnexcess-platform-sre"}]}