{"id":2008783,"url":"https://alion.io/job/nexcess-reliability-operations-specialist-2","title":"Reliability Operations Specialist","company":{"id":2155317,"name":"Nexcess","domain":"nexcess.com","url":"https://alion.io/company/nexcess-com","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Rippling","truth_index":{"grade":"B","score":75,"open_postings":4,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-08T05:49:30Z"}},"role":"Operations","role_family":"Operations","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Limassol, Cyprus"],"countries":["CY"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":29000,"max_usd":68000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":1732},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Commander.js","optional":false},{"name":"Incident Management","optional":false},{"name":"ITIL","optional":false},{"name":"ITSM","optional":false},{"name":"Linux","optional":false},{"name":"JavaScript","optional":true},{"name":"Node JS","optional":true}],"status":"live","first_seen_at":"2026-10-07T11:37:57Z","employer_posted_date":"2026-10-07","last_verified_at":"2026-10-08T21:31:14Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"About Nexcess\nNexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.\nThe Reliability Operations Specialist is responsible for driving operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, helping engineering teams improve service stability, reduce operational risk, and strengthen incident response processes.\nThe Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate incident response, oversee post-incident reviews, track corrective actions, and provide visibility into the health and reliability of the platform. This role does not have direct people management responsibilities but influences reliability outcomes across the organization through process ownership, collaboration, and data-driven decision making.\nResponsibilities\nIncident Management & Operational Excellence\nParticipate in major incident response activities and serve as an Incident Commander when assigned.\nFacilitate incident coordination, escalation, stakeholder communications, and status reporting during service-impacting events.\nSupport ongoing improvement of incident management processes, procedures, and operational readiness.\nDrive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).\nMaintain and apply the Criticality Matrix to tier services, infrastructure, and customer MRR impact. \nPost-Mortem Management & Corrective Actions\nCoordinate and facilitate post-mortem reviews following significant incidents.\nEnsure post-mortems are completed accurately, consistently, and within established timelines.\nSynthesize findings across incidents to identify trends, recurring issues, and systemic risks.\nMaintain accountability for corrective action tracking and closure.\nPromote a blameless culture of learning and continuous improvement and proactive/reactive problem management.\nReliability Strategy & Observability\nPartner with engineering teams to define and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across our product lines and services.\nSupport development and evolution of platform observability strategies, including monitoring, alerting, dashboards, and telemetry standards.\nAnalyze reliability metrics and operational trends to identify improvement opportunities.\nRecommend and track initiatives that improve platform stability, resiliency, and service performance\nPartner with DevOps to design JSM workflows for Incident, Problem, and Change processes while eliminating manual meetings through automation.\nChange Management & Governance\nEstablish change policy, lifecycle rules, risk assessments, CAB oversight, and approvals across standard, normal, and emergency changes.\nTrack and govern Change Failure Rates and change-related incidents to protect environment stability.\nService Transition & ITAM / Configuration Governance\nEmbed across Product Line pods to manage service acceptance, operational readiness, support models, and lifecycle status (supported/unsupported/EOL).\nMaintain asset governance, CMDB accuracy, Configuration Items (CIs), and service relationship mapping in JSM.\nReporting & Stakeholder Communication\nDevelop reliability reporting for engineering leadership and executive stakeholders.\nMaintain incident communication standards and stakeholder notification protocols.\nProvide regular reporting on reliability trends, corrective actions, incident performance, and service health.\nTranslate technical reliability metrics into actionable business insights.\nQualifications\nRequired\n3+ years of experience in Product Operations, Platform Operations, Technical Customer Support, Incident Coordination, or IT Service Management (ITSM/ITIL).\nExperience participating in or coordinating major incident response activities.\nKnowledge of incident management, root cause analysis, and post-mortem methodologies.\nExperience with monitoring, alerting, observability, or operational reporting tools.\nStrong analytical and organizational skills with attention to detail.\nExcellent written and verbal communication skills.\nAbility to work effectively across multiple teams and influence outcomes without direct authority.\nPreferred\nExperience working with SLOs, SLIs, and reliability metrics.\nFamiliarity with cloud infrastructure, Linux systems, networking, or distributed platforms.\nKnowledge of ITIL, operational excellence, or reliability engineering principles.\nExperience supporting high-availability SaaS, hosting, cloud, or infrastructure environments.\nExperience creating executive-level operational reports and presentations.","description_format":"text","description_chars":5157,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-10-07T12:53:06Z"}],"visa":[],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":0,"expected_fill_days":25,"reasons":["conf:0","velocity","win:early"],"computed_at":"2026-10-08T05:49:30Z"},"pay":null,"html_url":"https://alion.io/job/nexcess-reliability-operations-specialist-2","json_url":"https://alion.io/job/nexcess-reliability-operations-specialist-2.json","meta":{"generated_at":"2026-10-09T00:23:09Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":457,"day_limit":5000,"remaining_today":4543,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2155317},"rest":"https://alion.io/mcp/rest/get_company?id=2155317"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fnexcess-reliability-operations-specialist-2"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fnexcess-reliability-operations-specialist-2"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fnexcess-reliability-operations-specialist-2"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/nexcess-reliability-operations-specialist-2\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fnexcess-reliability-operations-specialist-2"}]}