807,812open jobs
51,801companies
127,727added this week
Browse all
Salary
≈ $74k – $183k per year (Estimated)
Location
In office (United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 25, 2026. Systems Engineering Solutions Corporation scores A on the Alion truth index.

Overview
Company
Impact
Profile match
SES is an industry leader in verification services with projects ranging from conformance with self-imposed sustainability standards to the functioning of national voluntary programs. Since 1998, SES has supported governmental and private clients ...

This role supports the U.S. Air Force Cloud One Architecture and Common Shared Services contract and currently has an opening for a Reliability Engineer. The Reliability Engineer is responsible for ensuring the availability, performance, scalability, and resiliency of mission-critical systems. This role applies software engineering principles to infrastructure and operations, with a strong emphasis on automation, monitoring, incident response, and continuous reliability improvement. The reliability engineer serves as the bridge between development, operations, and platform teams to ensure production systems consistently meet defined service level objectives (SLOs) while supporting rapid, safe delivery of new capabilities.

Location: This position will be hybrid remote. Candidates will be required to work onsite as needed. Candidates preferred to be located near Hanscom AFB (Boston, MA).

Requirements

System Reliability & Availability

  • Design, implement, and maintain highly available, fault-tolerant systems in cloud and hybrid environments
  • Define, measure, and report Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets
  • Identify reliability risks and implement mitigation strategies across the system lifecycle
  • Conduct capacity planning and performance modeling to ensure systems scale to meet demand

Monitoring, Observability & Alerting

  • Implement and manage monitoring, logging, and tracing solutions to provide full system observability
  • Define actionable alerting thresholds that minimize noise and enable rapid incident detection
  • Analyze trends and metrics to proactively identify potential reliability issues

Incident Response & Problem Management

  • Participate in on-call rotations and lead incident response activities for production systems
  • Coordinate troubleshooting efforts across development, infrastructure, and security teams
  • Conduct post-incident reviews (PIRs) and develop corrective and preventive action plans
  • Track recurring issues and ensure root causes are resolved

Automation & Engineering Excellence

  • Automate operational tasks to reduce manual intervention and operational risk
  • Develop scripts, tools, and services that improve system reliability and reduce mean time to recovery (MTTR)
  • Promote “automation over toil” and standardize operational workflows

Reliability-Focused Engineering

  • Participate in architecture and design reviews with an emphasis on reliability, resiliency, and recoverability
  • Validate disaster recovery (DR) and business continuity plans; test failover mechanisms
  • Support chaos engineering, fault injection testing, and resilience validation where appropriate

Collaboration & Governance

  • Partner with DevOps, Platform, and Security teams to ensure reliability aligns with delivery and compliance objectives
  • Document system reliability standards, runbooks, and operational procedures
  • Support compliance and audit activities (e.g., FedRAMP, FISMA, internal operational controls)

Required Skills:

· Bachelors and eight (8) years or more of experience; Masters and six (6) years or more of experience. Additional experience may be accepted in lieu of degree.

· Active Secret clearance at a minimum required to start

· US citizenship required

· Experience with cloud platforms (AWS, Azure, OCI, or GCP), including managed services

· Experience with containerized environments (Docker, Kubernetes)

· Familiarity with CI/CD pipelines and deployment automation

· SLOs and error budgets

· Capacity modeling and performance testing

· Strong understanding of:

· Distributed systems and high-availability architectures

· Linux/Windows system administration

· Networking fundamentals (DNS, TCP/IP, load balancing)

· Hands-on experience with:

· Monitoring and observability tools (e.g., Prometheus, Grafana, ELK/Elastic, Datadog, Azure Monitor)

· Infrastructure as Code (Terraform, ARM, CloudFormation)

· Scripting or programming languages (Python, Bash, Go, PowerShell, or similar)

· Experience supporting incident management and on-call operations

Preferred Skills

  • Experience with USAF Cloud One or Platform 1.
  • Experience with Zero Trust Architecture
  • Cloud certifications in AWS, Azure, Google, or Oracle clouds

Benefits

SES provides a competitive salary and the following benefits:

  • Medical
  • Dental
  • Vision
  • AD&D
  • STD
  • LTD
  • Company paid Life Insurance
  • 401k with employer contribution
  • Paid Time Off
  • Pet Insurance
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
807,812 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Industrial Engineering
Similar stack
Same company
United States
$80k – $120k per year • In office • 5+ years exp • Salt Lake City
Design
SolidWorks
Apply
$85k – $140k per year • In office • Salt Lake City
Cybersecurity
CAPA
Apply
$100k – $145k per year • In office • 5+ years exp • Bachelor's Degree • Salt Lake City
Design
SolidWorks
Apply
Quality Engineer 1 month ago
$60k – $80k per year • In office • 4+ years exp • Bachelor's Degree • Salt Lake City
Cybersecurity
CAPA
Management
Microsoft Office
Apply
$44k – $56k per year • In office • Internship • Bachelor's Degree • Salt Lake City
Design
SolidWorks
Management
Microsoft Office
Apply
In office • 12+ years exp • Raritan
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Vertex AI
Gemini
DevOps
Terraform
GCP
Azure
CI/CD
GitOps
Jenkins
AWS
Kubernetes
SRE
Platform Engineering
Service Mesh
Google GKE
Google Cloud Run
FinOps
IAM
DNS
Cybersecurity
Least Privilege
Apply
Technical Manager 1 day ago
≈ $131k – $262k per year (Estimated) • In office • Full-Time • Ashburn
Python
PowerShell
DevOps
Splunk
Ansible
Azure
AWS
Configuration Management
Nagios
TCP/IP
Management
ITIL
ITSM
Apply
In office • 10+ years exp • Master's Degree • Raritan
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Vertex AI
Fine-tuning
Scikit-learn
TensorFlow
PyTorch
Synthetic Data
Feature Store
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
AWS
Docker
Kubernetes
Google GKE
Google Cloud Run
IAM
Analytics
ETL/ELT
Apply
$69k – $103k per year • Equity 2–2% • In office • Full-Time • 3+ years exp • Paris
Python
TypeScript
AI/ML
Claude Code
AI Agents
OpenAI Codex
Apply
$48k – $62k per year • Remote (United Kingdom) • Full-Time
JavaScript
TypeScript
Databases
PostgreSQL
Frontend
React.js
Mobile
React Native
DevOps
GCP
CI/CD
AWS
Kubernetes
Trunk-Based Development
Management
Agile
Apply
≈ $93k – $200k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Orlando
DevOps
Nagios
TCP/IP
Wi-Fi
Management
ITIL
Apply
Technical Manager 1 day ago
≈ $131k – $262k per year (Estimated) • In office • Full-Time • Ashburn
Python
PowerShell
DevOps
Splunk
Ansible
Azure
AWS
Configuration Management
Nagios
TCP/IP
Management
ITIL
ITSM
Apply
≈ $79k – $179k per year (Estimated) • In office • Ashburn
DevOps
Datadog
Nginx
Cybersecurity
Okta
Ping Identity
Management
ServiceNow
ITIL
Apply
≈ $74k – $177k per year (Estimated) • In office • Ashburn
Management
ITIL
Apply
≈ $120k – $246k per year (Estimated) • In office • Secret • Full-Time • Bachelor's Degree • Boston
SQL
DevOps
Terraform
Ansible
Azure
CI/CD
Windows Server
Jenkins
Git
AWS
Kubernetes
Bitbucket
Cybersecurity
PCI DSS
SOC 2
HIPAA
Zero Trust
Defense in Depth
Active Directory
Management
Confluence
Jira
SharePoint
Apply
≈ $54k – $105k per year (Estimated) • In office • Full-Time • 3+ years exp • High School Diploma • Cedar Rapids
Management
Microsoft Office
Apply
Quality Engineer 1 day ago
≈ $55k – $107k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Atlanta
Apply
$91k – $114k per year • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree • United States
Design
AutoCAD
Management
Microsoft Office
Apply
≈ $136k – $265k per year (Estimated) • Equity • In office • Full-Time • 5+ years exp • United States
SQL
Databases
Snowflake
Google BigQuery
BigQuery
AI/ML
Claude
Claude Code
Model Context Protocol
Prompt Engineering
AI Agents
DevOps
SLI/SLO/SLA
Robotics
Apollo
Management
n8n
Apply
Machine Operator I 1 day ago
$38k – $44k per year • In office • Full-Time • 1+ year exp • High School Diploma • United States
Apply
See all jobs
This is one of many
807,812 more open roles from verified company boards, updated every day.