1,437,834open jobs
84,810companies
219,155added this week
Browse all
Salary
≈ $12k – $30k per year (Estimated)
Location
In office (India)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 7, 2026. Nexcess scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Nexcess is a managed cloud hosting provider with built-in compliance and predictable costs.

About Nexcess

Nexcess provides specialty cloud solutions for organizations where performance and compliance have to coexist. We serve businesses worldwide, from agencies scaling client sites to enterprises running mission-critical operations. We've built our reputation on deep technical expertise and genuine partnership with every client we work with. Behind every environment we manage is a team of people who take the craft seriously and keep showing up when it matters.

The Platform SRE is a member of the Platform SRE team, the group responsible for the operational health, resiliency, and modernization of our hosting and cloud infrastructure. This role sits at the intersection of systems engineering, site reliability engineering, and automation : combining hands-on technical debt remediation with large-scale automation and reliability practice.

The Platform SRE team owns three core mandates:

1. Technical debt reduction : systematically identifying and remediating aging firmware, kernels, operating systems, and software across the managed hosting fleet and managed application environments.

2. Automation of critical operational workflows : provisioning, patching, remediation, and release processes across both managed apps and managed hosting fleets, with automated release planning and execution under human supervision (not “automation for automation’s sake,” but automation with a human checkpoint before production impact).

3. Reliability and incident support : defining, instrumenting, and tracking SLIs/SLOs for platform engineering and operations, visualized through dashboards and reporting tools, and providing fast, expert frontline response and remediation during service-impacting events (incident command and process ownership sit with the Incident Management team; this team is the technical responder, not the incident owner).

This position serves as a senior technical point of contact for platform engineering, driving initiatives around scalability, fault tolerance, automation, and operational excellence across production infrastructure.

Key Responsibilities

Technical Debt & Platform Modernization

  • Own the lifecycle of firmware, kernel, OS, and software patching across the managed hosting fleet and managed application environments
  • Build a standing inventory and risk model of technical debt (end-of-life OS versions, unpatched firmware, deprecated software) and drive prioritized remediation plans
  • Evaluate and implement infrastructure modernization initiatives, replacing manual or legacy processes with supportable, automated alternatives

Automation & Release Engineering

  • Design and build automation for provisioning, deployment, patching, remediation, and configuration management across managed apps and managed hosting fleets
  • Own the design of automated release pipelines : planning, staging, and executing releases with defined human-in-the-loop approval gates
  • Develop self-healing and auto-remediation capability for common failure modes to reduce manual operational load
  • Support and extend CI/CD workflows and infrastructure-as-code practices across the platform

Reliability Engineering, SLIs/SLOs & Observability

  • Define SLIs and SLOs for platform engineering and operations in partnership with engineering and product stakeholders
  • Instrument systems to measure SLIs accurately and build SLO tracking into standard reporting
  • Build and maintain dashboards (e.g., Grafana, Datadog, or equivalent visualization tooling) to make SLI/SLO performance, error budgets, and platform health visible to engineering and leadership
  • Continuously improve platform observability : monitoring, alerting, logging, and tracing : across distributed and containerized environments

Incident Response & Remediation (Support Role)

  • Serve as the frontline technical responder: acknowledge pages quickly, diagnose, and remediate platform-level issues
  • Partner with the Incident Management team, who own incident command, severity classification, and customer communication : this role provides the technical hands and expertise, not incident ownership
  • Contribute technical findings to blameless root cause analysis (RCA) and own follow-through on corrective actions for platform systems
  • Maintain runbooks and on-call readiness for platform and infrastructure systems
  • Track incident trends on platform systems and feed them back into the technical debt and automation roadmap

Collaboration & Technical Leadership

  • Partner with software engineering teams on platform architecture, operational readiness reviews, and scalability initiatives
  • Support platform security, compliance, and operational governance requirements
  • Mentor engineers and contribute to technical leadership and knowledge-sharing across the team
  • Maintain clear operational documentation and contribute to team standards and process improvement
  • Other duties as assigned

Requirements

  • 3-5+ years of experience in platform engineering, systems engineering, SRE, or infrastructure operations (level based on experience and scope)
  • Advanced Linux systems administration and troubleshooting expertise, including kernel and firmware-level familiarity
  • Strong experience with Kubernetes, Docker, and container orchestration/distributed systems
  • Hands-on automation and infrastructure-as-code experience (e.g., Terraform, Ansible, Puppet/Chef, or equivalent)
  • Experience building or maintaining CI/CD and automated release/deployment pipelines
  • Experience defining and tracking SLIs/SLOs and working with observability/visualization tools (e.g., Grafana, Datadog, Prometheus, or equivalent)
  • Experience supporting enterprise-scale, high-concurrency, or customer-impacting production environments
  • Demonstrated experience as a technical responder in production incidents, including root cause analysis and corrective action follow-through
  • Strong scripting ability (e.g., Python, Bash, Go) for automation and tooling
  • Strong troubleshooting skills across compute, network, storage, and application layers
  • Experience supporting cloud-hosted, managed hosting, or hybrid infrastructure environments
  • Ability to lead technical initiatives, mentor others, and communicate clearly across teams

Preferred Qualifications

  • Experience owning fleet-wide firmware/OS patch management programs at scale
  • Experience designing human-in-the-loop release automation or progressive delivery systems (canary, blue/green)
  • Familiarity with error budgets and SLO-driven prioritization frameworks
  • Experience with configuration/patch management at scale across heterogeneous hardware fleets
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,437,834 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
India
≈ $16k – $35k per year (Estimated) • In office • 12+ years exp • Hyderabad
Databases
Databricks
Delta Lake
Microsoft Fabric
DevOps
Azure DevOps
Azure
CI/CD
Platform Engineering
Bicep
Apply
≈ $12k – $31k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Node JS
Node JS
Commander.js
Databases
PostgreSQL
AI/ML
AI Agents
LLM
DevOps
Terraform
GitHub Actions
OpenTelemetry
PagerDuty
CI/CD
AWS
Kubernetes
Platform Engineering
Self-Healing
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
Progressive Delivery
SLI/SLO/SLA
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
DNS
Cybersecurity
ISO 27001
SOC 2
Apply
$17k per year • In office • Full-Time • 4+ years exp • Bengaluru
SQL
Databases
Snowflake
Db2
DevOps
AWS
Analytics
ETL/ELT
Apply
≈ $11k – $27k per year (Estimated) • In office • Full-Time • 4+ years exp • Pune
SQL
DevOps
Azure
IAM
API Gateway
Cybersecurity
Okta
ISO 27001
GDPR
Least Privilege
Ping Identity
Management
ITIL
Apply
≈ $9.5k – $24k per year (Estimated) • In office • Full-Time • 3+ years exp • Mumbai
JavaScript
PHP
PHP
WordPress
Databases
MySQL
DevOps
Azure
DNS
Wi-Fi
Cybersecurity
ISO 27001
Management
Zapier
Outlook
SharePoint
Apply
Nutanix Engineer 2 days ago
≈ $63k – $161k per year (Estimated) • In office • 5+ years exp • Pago Pago
Python
PowerShell
DevOps
Terraform
Ansible
GCP
VMWare
Azure
AWS
Hyper-V
TCP/IP
DNS
DHCP
Apply
≈ $13k – $31k per year (Estimated) • In office • 5+ years exp • Pune
Python
PowerShell
DevOps
Terraform
Ansible
GCP
VMWare
Azure
AWS
Hyper-V
TCP/IP
DNS
DHCP
Apply
≈ $63k – $115k per year (Estimated) • In office • 6+ years exp • Amsterdam
Python
SQL
PowerShell
Bash
Python
SQLAlchemy
FastAPI
Databases
Databricks
Apache Kafka
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Docker
Kubernetes
Platform Engineering
Linux
Management
Agile
QA
Pytest
Apply
≈ $47k – $119k per year (Estimated) • In office • 5+ years exp • Cavan
Python
PowerShell
DevOps
Terraform
Ansible
GCP
VMWare
Azure
AWS
Hyper-V
TCP/IP
DNS
DHCP
Apply
≈ $53k – $134k per year (Estimated) • In office • 5+ years exp • Canberra
Python
PowerShell
DevOps
GCP
VMWare
Azure
AWS
Hyper-V
TCP/IP
DNS
DHCP
Apply
≈ $42k – $106k per year (Estimated) • In office • Full-Time • 5+ years exp • Limassol
AI/ML
AI Agents
DevOps
VMWare
Linux
Apply
≈ $42k – $106k per year (Estimated) • In office • Full-Time • 5+ years exp • Bratislava
AI/ML
AI Agents
DevOps
Puppet
Ansible
Configuration Management
Linux
VPN
Apply
DevOps 4 days ago
≈ $42k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Limassol
Go
PHP
Perl
AI/ML
Model Context Protocol
DevOps
CI/CD
GitHub
Apply
≈ $42k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Bratislava
AI/ML
AI Agents
DevOps
Puppet
Ansible
Configuration Management
KVM
Linux
Apply
≈ $42k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Limassol
AI/ML
AI Agents
DevOps
Puppet
Ansible
Configuration Management
Linux
Apply
Cloud Architect 1 day ago
≈ $22k – $51k per year (Estimated) • In office • Bachelor's Degree • India
SQL
DevOps
Rest API
Terraform
CloudFormation
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
≈ $17k – $37k per year (Estimated) • In office • Full-Time • 6+ years exp • India
Python
JavaScript
TypeScript
SQL
PowerShell
C#
Groovy
C#
.NET
Databases
MySQL
Redis
ElasticSearch
Frontend
Angular
npm
DevOps
Terraform
Azure DevOps
New Relic
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Cloudflare
JFrog Artifactory
AWS Lambda
Amazon S3
Windows
TCP/IP
DNS
Cybersecurity
SonarQube
SOC 2
HIPAA
Management
Jira
Agile
QA
Sentry
Apply
In office • India
DevOps
Rest API
SOAP
Apply
Hardware Architect 1 day ago
≈ $26k – $52k per year (Estimated) • In office • 15+ years exp • India
Apply
≈ $15k – $37k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • India
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
AI/ML
Time Series Forecasting
Frontend
Angular
DevOps
Azure DevOps
Prometheus
WebSockets
Azure
CI/CD
Git
Grafana
Cybersecurity
SonarQube
ISO 27001
GDPR
IoT
MQTT
Management
Agile
Apply
See all jobs
This is one of many
1,437,834 more open roles from verified company boards, updated every day.