1,401,056open jobs
81,239companies
208,212added this week
Browse all
Salary
$85k – $100k per year
Location
Remote (United States, India, United Kingdom, Bulgaria)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Jun 9, 2026. Nexcess scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Nexcess is a managed cloud hosting provider with built-in compliance and predictable costs.

About Nexcess

Nexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.

- Incident, Problem & Change Management, Reporting & Insights -

About the rol e

We're looking for a Reliability Operations Specialist to help drive operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, helping teams improve service stability, reduce operational risk, and strengthen operational readiness across the organization.

The Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate incident response, oversee post-incident reviews, track corrective actions, and provide visibility into the health and reliability of our platforms. This role does not have direct people management responsibilities but plays a critical role in influencing reliability outcomes through collaboration, process ownership, and data-driven decision making.

Location: Remote

Employment type: Permanent, Full-time

Pay Range: $85,000 - $100,000 Annually, The final compensation offered will be determined based on factors including location, experience, skills, qualifications, and market conditions.

What you'll do

Incident Management & Operational Excellence

  • Participate in major incident response activities and serve as an Incident Commander when assigned.
  • Facilitate incident coordination, escalation, stakeholder communications, and status reporting during service-impacting events.
  • Support ongoing improvement of incident management processes, procedures, and operational readiness.
  • Drive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Maintain and apply the Criticality Matrix to tier services, infrastructure, and customer MRR impact.

Mortem Management & Corrective Actions

  • Coordinate and facilitate post-mortem reviews following significant incidents.
  • Ensure post-mortems are completed accurately, consistently, and within established timelines.
  • Synthesize findings across incidents to identify trends, recurring issues, and systemic risks.
  • Maintain accountability for corrective action tracking and closure.
  • Promote a blameless culture of learning and continuous improvement and proactive/reactive problem management.

Reliability Strategy & Observability

  • Partner with engineering teams to define and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across our product lines and services.
  • Support development and evolution of platform observability strategies, including monitoring, alerting, dashboards, and telemetry standards.
  • Analyze reliability metrics and operational trends to identify improvement opportunities.
  • Recommend and track initiatives that improve platform stability, resiliency, and service performance
  • Partner with DevOps to design JSM workflows for Incident, Problem, and Change processes while eliminating manual meetings through automation.

Change Management & Governance

  • Establish change policy, lifecycle rules, risk assessments, CAB oversight, and approvals across standard, normal, and emergency changes.
  • Track and govern Change Failure Rates and change-related incidents to protect environment stability.

Service Transition & ITAM / Configuration Governance

  • Embed across Product Line pods to manage service acceptance, operational readiness, support models, and lifecycle status (supported/unsupported/EOL).
  • Maintain asset governance, CMDB accuracy, Configuration Items (CIs), and service relationship mapping in JSM

Reporting & Stakeholder Communication

  • Develop reliability reporting for engineering leadership and executive stakeholders.
  • Maintain incident communication standards and stakeholder notification protocols.
  • Provide regular reporting on reliability trends, corrective actions, incident performance, and service health.
  • Translate technical reliability metrics into actionable business insights.

What you bring

  • 3+ years of experience in Product Operations, Platform Operations, Technical Customer Support, Incident Coordination, Site Operations, IT Service Management (ITSM), or a related operational role.
  • Experience participating in or coordinating major incident response activities.
  • Knowledge of incident management, root cause analysis, problem management, and post-mortem methodologies.
  • Experience working with monitoring, alerting, observability, or operational reporting tools.
  • Strong analytical and organizational skills with exceptional attention to detail.
  • Excellent written and verbal communication skills.
  • Ability to work effectively across multiple teams and influence outcomes without direct authority.
  • Strong problem-solving skills and the ability to remain calm and organized during high-pressure situations
  • Hands-on proficiency with Jira Service Management (JSM) workflows, automation, and data quality governance.

Preferred Qualifications

  • Experience working with Service Level Objectives (SLOs), Service Level Indicators (SLIs), and reliability metrics.
  • Familiarity with Linux systems, cloud infrastructure, networking concepts, hosting platforms, or distributed systems.
  • Knowledge of ITIL, operational excellence frameworks, or Site Reliability Engineering (SRE) principles.
  • Experience supporting high-availability SaaS, hosting, cloud, or infrastructure environments.
  • Experience creating executive-level operational reports, dashboards, and presentations.
  • Experience using observability and incident management platforms.

What We Offer

  • Comprehensive benefits package
  • Traditional and Roth 401(k) with company matching
  • A collaborative, team-oriented culture
  • Consistent and predictable work hours
  • Engaging, varied work that keeps each day different
  • Opportunities to contribute ideas and influence how work gets done

Disclaimer:

This job description is only a summary of the typical functions of the position. It is not intended to be an exhaustive or comprehensive list of all job responsibilities, tasks, or duties. Additional duties and tasks may be assigned as part of the job function. Nexcess reserves the right to modify, interpret, or apply this job description in a way that best supports the organizational needs. The job description in no way creates or implies an employment contract. The employment contract remains “at will”.

Equal Employment Opportunity Policy:

Nexcess is committed to offering equal employment opportunity without regard to age, color, disability, gender, gender identity, genetic information, marital status, military status, national origin, race, religion, sexual orientation, veteran status, or any other legally protected characteristic.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,401,056 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Operations
Similar stack
Same company
In your city
≈ $14k – $30k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Coimbatore
Analytics
Microsoft Excel
Management
Outlook
Microsoft Office
Apply
≈ $8k – $20k per year (Estimated) • In office • Contractor • 2+ years exp • Bachelor's Degree • Hyderabad
Management
Microsoft Office
Apply
≈ $22k – $58k per year (Estimated) • Remote (India) • Full-Time • 5+ years exp • Bengaluru
Management
Agile
Apply
≈ $12k – $28k per year (Estimated) • Remote (India) • Full-Time • Bengaluru
Management
Agile
Microsoft Office
Apply
≈ $43k – $98k per year (Estimated) • Remote (Spain) • 5+ years exp • Bachelor's Degree • Barcelona
Analytics
Power BI
Microsoft Excel
Apply
≈ $33k – $56k per year (Estimated) • In office • Internship • Germany
Python
Java
Management
Confluence
Jira
Microsoft Office
Apply
In office • 5+ years exp • Bachelor's Degree
AI/ML
Copilot
DevOps
Azure
Incident Management
Cybersecurity
Active Directory
Analytics
Power BI
Management
SharePoint
Apply
≈ $33k – $57k per year (Estimated) • In office • Internship • Germany
Management
Confluence
Jira
Microsoft Office
Apply
≈ $34k – $59k per year (Estimated) • In office • Internship • Germany
Python
Management
Jira
Microsoft Office
Apply
≈ $27k – $59k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Hefei
Management
Confluence
Jira
Apply
$165k – $192k per year • In office • Full-Time • 8+ years exp • Guildford
Apply
$110k – $130k per year • Remote (United States) • Full-Time • 4+ years exp • Bachelor's Degree
DevOps
GCP
Azure
AWS
Incident Management
Linux
DNS
Cybersecurity
ISO 27001
SOC 2
HIPAA
Management
ITIL
Apply
≈ $91k – $173k per year (Estimated) • Remote (United States) • Full-Time • 7+ years exp
PHP
Perl
Databases
PostgreSQL
MariaDB
AI/ML
AI Agents
Apply
≈ $135k – $226k per year (Estimated) • Remote (United States) • Full-Time • 7+ years exp
PHP
Perl
Databases
PostgreSQL
MariaDB
AI/ML
AI Agents
Apply
Director, Engineering 18 days ago
≈ $154k – $276k per year (Estimated) • Remote (United States) • Full-Time • 12+ years exp
PHP
Perl
PHP
WordPress
WooCommerce
Magento
AI/ML
AI Agents
DevOps
VMWare
Trunk-Based Development
API Gateway
Linux
Management
Stripe
Apply
See all jobs
This is one of many
1,401,056 more open roles from verified company boards, updated every day.