372,254open jobs
9,642companies
49,764added this week
Browse all
Salary
$113k – $220k per year (Estimated)
Location
Remote (United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Everbridge is a critical event management software company headquartered in Burlington, Massachusetts, and founded in 2002. The company runs a platform that detects threats such as severe weather, active shooters, or IT outages, then locates affected people and sends coordinated alerts across phone, SMS, email, and public warning channels. It serves corporations, hospitals, universities, and government agencies including national public warning systems, and was taken private by Thoma Bravo in 2024.

At Everbridge, reliability isn’t just about uptime-it’s about ensuring that critical systems are available when they matter most. Every improvement you make helps organizations deliver life-saving communications and maintain operations during emergencies.

As a Senior Site Reliability Engineer II, you’ll do more than operate infrastructure. You’ll improve the resiliency of our engineering organization by building reliable platforms, eliminating operational toil, mentoring engineers, and helping teams design systems that are secure, scalable, and resilient by

default.

This is a highly collaborative technical leadership role for an engineer who enjoys solving systemic problems, influencing architecture, and enabling others to build reliable software.

What you'll do:

  • Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently.
  • Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, networking, and developer platforms.
  • Influence engineering decisions through technical expertise, collaboration, and data.
  • Help engineering teams become increasingly self-sufficient through coaching, automation, and well-designed platform capabilities.
  • Continuously improve operational excellence by reducing complexity, eliminating manual work, and strengthening engineering practices.
  • Design and implement solutions that improve the availability, scalability, performance, and resilience of our platform.
  • Build automation that eliminates repetitive operational work and reduces engineering toil.
  • Improve observability, monitoring, alerting, and operational readiness across the organization.
  • Use production data, reliability metrics, and engineering judgment to identify systemic improvements.
  • Lead complex cross-functional engineering initiatives from design through production.
  • Partner with architects, software engineers, security, product, and platform teams to build resilient systems from the beginning.
  • Review architectures and designs with a focus on reliability, scalability, recoverability, and operational excellence.
  • Establish and evolve engineering standards, best practices, and operational readiness guidance.
  • Partner directly with engineering teams to improve the reliability of the services they own.
  • Coach teams on observability, incident response, disaster recovery, capacity planning, and production readiness.
  • Help teams adopt SLOs, error budgets, meaningful alerting, and engineering practices that improve customer outcomes.
  • Make the right engineering decisions easier through automation, paved roads, and self-service capabilities.
  • Participate in an on-call rotation supporting critical production systems.
  • Lead the technical response during high-severity incidents.
  • Facilitate blameless post-incident reviews that focus on learning and systemic improvement.
  • Drive corrective actions through completion and measure their effectiveness over time.
  • Share knowledge through documentation, technical design reviews, and collaborative problem solving.
  • Foster a culture of ownership, continuous improvement, operational excellence, and customer focus.

What you'll bring:

  • Experience with:
  • Designing and operating complex production systems.
  • Cloud infrastructure and cloud-native architectures.
  • Distributed systems and container platforms.
  • Infrastructure as Code and automation.
  • CI/CD and software delivery practices.
  • Observability, monitoring, logging, and telemetry.
  • Incident response and operational excellence.
  • Reliability engineering principles including SLOs, SLIs, capacity planning, and performance optimization.
  • Writing software or automation using one or more modern programming languages.
  • Linux and networking fundamentals.
  • Experience working within regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001 is a plus.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
372,254 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$136k – $275k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • United States
Go
JavaScript
TypeScript
Databases
DynamoDB
AI/ML
AI Agents
Function Calling
Model Context Protocol
Prompt Engineering
RAG
DevOps
Amazon EC2
Amazon ECS
Amazon EKS
Amazon S3
AWS
AWS Fargate
AWS Lambda
AWS Step Functions
CI/CD
CloudFormation
Docker
GitHub
Kubernetes
Terraform
Apply
$20k – $51k per year (Estimated) • In office • Full-Time • 6+ years exp • India
C#
SQL
TypeScript
JavaScript
C#
ASP.NET Core
Entity Framework Core
xUnit
Databases
Apache Kafka
RabbitMQ
Frontend
Angular
Mobile
Dependency Injection
DevOps
Azure
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Rest API
Apply
$115k – $226k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • United States
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
Anomaly Detection
DevOps
Amazon CloudWatch
Amazon S3
AWS
CI/CD
Git
Apply
$20k – $49k per year (Estimated) • In office • Full-Time • 5+ years exp • Chennai
Apex
Apex
Lightning Web Components
AI/ML
AI Agents
Agentforce
DevOps
CI/CD
GitHub
Management
Jira
ServiceNow
Marketing
Salesforce
Apply
$151k – $301k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United States
AI/ML
AI Agents
DevOps
AWS
Kubernetes
OpenShift
HPC
Apply
$66k – $110k per year • Remote • Full-Time • 3+ years exp • Bachelor's Degree • Auckland
Management
Confluence
Apply
$60k – $144k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree
JavaScript
Python
TypeScript
Java
Java
Spring Boot
Databases
Apache Kafka
MariaDB
PostgreSQL
AI/ML
Claude
Claude Code
Copilot
Cursor
OpenAI Codex
Frontend
Angular
DevOps
Ansible
Docker
Helm
Kubernetes
GitHub
Apply
$27k – $67k per year (Estimated) • Remote • Full-Time • 5+ years exp
Java
JavaScript
Java
Spring Boot
AI/ML
Claude
Claude Code
Copilot
Cursor
Ignite
Windsurf
PyTorch
OpenAI Codex
Frontend
React.js
DevOps
AWS
CI/CD
Docker
Incident Management
Kubernetes
Platform Engineering
GitHub
Apply
$27k – $66k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
C#
JavaScript
TypeScript
C#
.NET
Databases
Apache Kafka
Frontend
npm
React.js
DevOps
AWS
CI/CD
Docker
Kubernetes
GitLab
Apply
Remote • Full-Time
Databases
PostgreSQL
DevOps
Incident Management
Platform Engineering
Apply
See all jobs
This is one of many
372,254 more open roles from verified company boards, updated every day.