1,348,834open jobs
78,754companies
206,738added this week
Browse all
Salary
≈ $29k – $68k per year (Estimated)
Location
In office (Limassol)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Oct 7, 2026. Nexcess scores B on the Alion truth index.

Overview
Company
Impact
Profile match
We offer expert-managed cloud infrastructure with built-in compliance and predictable costs. Run mission-critical workloads without infrastructure overhead.

About Nexcess

Nexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.

The Reliability Operations Specialist is responsible for driving operational excellence across incident management, service reliability, observability, and continuous improvement initiatives. This role serves as a central coordinator and subject matter expert for reliability practices, helping engineering teams improve service stability, reduce operational risk, and strengthen incident response processes.

The Reliability Operations Specialist partners closely with engineering, infrastructure, security, and operations teams to facilitate incident response, oversee post-incident reviews, track corrective actions, and provide visibility into the health and reliability of the platform. This role does not have direct people management responsibilities but influences reliability outcomes across the organization through process ownership, collaboration, and data-driven decision making.

Responsibilities

Incident Management & Operational Excellence

  • Participate in major incident response activities and serve as an Incident Commander when assigned.
  • Facilitate incident coordination, escalation, stakeholder communications, and status reporting during service-impacting events.
  • Support ongoing improvement of incident management processes, procedures, and operational readiness.
  • Drive initiatives focused on reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR).
  • Maintain and apply the Criticality Matrix to tier services, infrastructure, and customer MRR impact.

Post-Mortem Management & Corrective Actions

  • Coordinate and facilitate post-mortem reviews following significant incidents.
  • Ensure post-mortems are completed accurately, consistently, and within established timelines.
  • Synthesize findings across incidents to identify trends, recurring issues, and systemic risks.
  • Maintain accountability for corrective action tracking and closure.
  • Promote a blameless culture of learning and continuous improvement and proactive/reactive problem management.

Reliability Strategy & Observability

  • Partner with engineering teams to define and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) across our product lines and services.
  • Support development and evolution of platform observability strategies, including monitoring, alerting, dashboards, and telemetry standards.
  • Analyze reliability metrics and operational trends to identify improvement opportunities.
  • Recommend and track initiatives that improve platform stability, resiliency, and service performance
  • Partner with DevOps to design JSM workflows for Incident, Problem, and Change processes while eliminating manual meetings through automation.

Change Management & Governance

  • Establish change policy, lifecycle rules, risk assessments, CAB oversight, and approvals across standard, normal, and emergency changes.
  • Track and govern Change Failure Rates and change-related incidents to protect environment stability.

Service Transition & ITAM / Configuration Governance

  • Embed across Product Line pods to manage service acceptance, operational readiness, support models, and lifecycle status (supported/unsupported/EOL).
  • Maintain asset governance, CMDB accuracy, Configuration Items (CIs), and service relationship mapping in JSM.

Reporting & Stakeholder Communication

  • Develop reliability reporting for engineering leadership and executive stakeholders.
  • Maintain incident communication standards and stakeholder notification protocols.
  • Provide regular reporting on reliability trends, corrective actions, incident performance, and service health.
  • Translate technical reliability metrics into actionable business insights.

Qualifications

Required

  • 3+ years of experience in Product Operations, Platform Operations, Technical Customer Support, Incident Coordination, or IT Service Management (ITSM/ITIL).
  • Experience participating in or coordinating major incident response activities.
  • Knowledge of incident management, root cause analysis, and post-mortem methodologies.
  • Experience with monitoring, alerting, observability, or operational reporting tools.
  • Strong analytical and organizational skills with attention to detail.
  • Excellent written and verbal communication skills.
  • Ability to work effectively across multiple teams and influence outcomes without direct authority.

Preferred

  • Experience working with SLOs, SLIs, and reliability metrics.
  • Familiarity with cloud infrastructure, Linux systems, networking, or distributed platforms.
  • Knowledge of ITIL, operational excellence, or reliability engineering principles.
  • Experience supporting high-availability SaaS, hosting, cloud, or infrastructure environments.
  • Experience creating executive-level operational reports and presentations.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,348,834 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Operations
Similar stack
Same company
Limassol
≈ $22k – $50k per year (Estimated) • In office • 16+ years exp • Chennai
Apply
$60k – $66k per year • In office • Full-Time • Brisbane
Management
Microsoft Office
Apply
≈ $76k – $164k per year (Estimated) • In office • 10+ years exp • Seoul
Apply
In office • Bülach
Apply
≈ $7k – $18k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Sānand
Analytics
Power BI
Microsoft Excel
Apply
≈ $42k – $90k per year (Estimated) • Hybrid • Full-Time • Rennes
DevOps
Rest API
Azure
Linux
Windows
Management
ITSM
Apply
≈ $59k – $149k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Santa Venera
Python
Bash
Databases
PostgreSQL
Redis
OpenSearch
DevOps
Ansible
Prometheus
Windows Server
Git
AWS
Ubuntu
Grafana
Configuration Management
Amazon EC2
Amazon S3
IAM
Amazon CloudWatch
Linux
Windows
Apply
≈ $76k – $158k per year (Estimated) • In office • Bachelor's Degree • Durham
SQL
ABAP
ABAP
SAP Fiori
CDS Views
Web Dynpro ABAP
Databases
SAP HANA
DevOps
Linux
Unix
Apply
≈ $14k – $26k per year (Estimated) • Hybrid • 2+ years exp • Saint Petersburg
Python
Bash
Databases
InfluxDB
DevOps
Zabbix
Debian
Prometheus
GitLab CI
CI/CD
Docker
Ubuntu
Grafana
Proxmox VE
KVM
GitLab
Linux
TCP/IP
VLAN
Apply
≈ $28k – $73k per year (Estimated) • In office • Contractor • France
Python
Java
PHP
SQL
Databases
SAP BW
Management
ITIL
Apply
≈ $64k – $115k per year (Estimated) • In office • Full-Time • 8+ years exp • Guildford
Apply
$85k – $100k per year • Remote (United States, India, United Kingdom, Bulgaria) • Full-Time • 3+ years exp
JavaScript
Node JS
Node JS
Commander.js
DevOps
Incident Management
Linux
Management
Jira
ITIL
ITSM
Apply
≈ $43k – $119k per year (Estimated) • In office • Full-Time • 5+ years exp • Sofia
SQL
Perl
Databases
PostgreSQL
AI/ML
AI Agents
DevOps
VMWare
Apply
≈ $42k – $106k per year (Estimated) • In office • Full-Time • 5+ years exp • Limassol
AI/ML
AI Agents
DevOps
VMWare
Linux
Apply
In office • Full-Time • 5+ years exp • Limassol
JavaScript
TypeScript
Databases
PostgreSQL
MariaDB
AI/ML
AI Agents
Frontend
Vue.js
React.js
DevOps
VPN
Apply
In office • Full-Time • Limassol
Apply
Architect 2 days ago
≈ $51k – $110k per year (Estimated) • In office • Full-Time • 3+ years exp • Limassol
Design
AutoCAD
Apply
Project Director 2 days ago
≈ $63k – $141k per year (Estimated) • In office • Full-Time • Limassol
Apply
≈ $32k – $83k per year (Estimated) • In office • Full-Time • Limassol
Management
Google Workspace
Microsoft Office
Apply
In office • Full-Time • Limassol
JavaScript
TypeScript
Frontend
Zustand
Redux
Three.JS
React.js
PixiJS
Jotai
React Three Fiber
Mobile
State Management
DevOps
WebRTC
WebSockets
Game Dev
GLSL
Design
Blender
Rive
QA
Jest
Vitest
Apply
See all jobs
This is one of many
1,348,834 more open roles from verified company boards, updated every day.