824,589open jobs
53,146companies
135,057added this week
Browse all
Salary
$105k – $160k per year
Location
In office (Seattle)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Sep 15, 2026.

Overview
Company
Impact
Profile match
Amazon is an American technology and retail conglomerate founded by Jeff Bezos in 1994 as an online bookstore and headquartered in Seattle, Washington. It operates the world's largest online marketplace together with a global logistics network, physical grocery stores and a third-party seller platform that accounts for most units sold. Amazon Web Services, launched in 2006, is the leading public cloud provider and generates the majority of the group's operating profit, while advertising, Prime Video, Alexa devices and Kuiper satellite broadband round out the business.

As a Systems Engineer on AWS Incident Response, you will be on the front line of AWS incident response. You will lead high-severity calls, triage complex failures across distributed systems, coordinate resolver teams, and drive incidents to mitigation in real time while millions of customers depend on the outcome. Between incidents you will obsess over metrics and detection analysis, building dashboards and mechanisms that surface problems before customers notice, and you will drive operational improvements that make the incident management ecosystem faster and more accurate.

You will make real-time decisions under pressure, deep-diving the largest and most complex technical environment in the world. You will develop expertise across AWS services, networking, and infrastructure, building a breadth of knowledge that few roles offer. Your scope spans all of AWS rather than a single service. You will own operational processes end to end and use data to find the next improvement in how we detect and mitigate faster. You will also have the opportunity to grow your development skills by taking on coding projects that accelerate incident response and reduce toil.

This role includes participation in an on-call rotation covering weekdays, weekends, and holidays. On-call shifts fall within your local daytime hours.

Key job responsibilities

- Lead high-severity incident response calls end to end: assess impact, coordinate resolvers across AWS service teams, communicate clearly under pressure, manage escalations, and drive the incident to mitigation with documentation throughout.

- Own and run operational health reviews, and build and maintain the dashboards, metrics, and monitoring that surface trends before they become incidents.

- Improve detection accuracy and speed. Identify patterns across events and build proactive mechanisms that prevent recurrence.

- Deep-dive operational data to find systemic issues, measure response effectiveness, and prioritize improvements against what the data shows is degrading.

- Identify gaps in operational processes, documentation, and tooling, and build or improve mechanisms that reduce time-to-detection and time-to-mitigation.

- Apply scripting, automation, and generative AI to accelerate incident response and reduce toil, including where AI can augment human judgment during an incident or surface insight from operational data at scale.

- Work with service teams so that learnings from each incident drive corrective actions to completion, closing the loop between what broke and what gets fixed.

- Mentor peers in your areas of technical and operational strength.

A day in the life

When you are on call, incidents take priority. You join the incident bridge, assess the scope of impact from real-time metrics and dashboards, engage the resolver teams that own the affected services, and drive the event to mitigation, escalating when progress stalls. Off-call, you work to reduce time-to-detection and time-to-mitigation: deep-diving recent events for detection that lagged the impact or alarms that trigger without customer impact, correcting the signal behind them, leading operational health reviews, driving other teams' corrective actions to completion, and owning projects that strengthen AIR's detection and incident management, including automation you can build yourself.

About the team

AWS Incident Response (AIR), part of the AWS Resilience organization, ensures high availability of AWS. When major incidents hit, AIR leads the response, coordinating resolver teams across AWS and driving mitigation. We move fast, but not carelessly, obsessing over observability of the cloud and continuously improving our detection and response speed and accuracy. Each incident feeds improvements to our detection, processes, and tooling, making the next one shorter or preventing it entirely. This is a high-visibility, high-impact role with a global view of AWS health that few teams get to see.

Basic qualifications

- 3+ years of systems engineering, or 3+ years of technical support experience

- Experience in written and verbal communication skills to communicate with technical and non-technical audiences, including senior leadership

- Experience scripting in one or more language (e.g. Bash, Python, Perl, Ruby), or experience scripting in modern programming languages

- Understanding of operating systems (Linux), networking fundamentals, and distributed systems

- Experience with operational monitoring, alerting, and metrics (CloudWatch, Datadog, Grafana, or equivalent)

- Demonstrated ability to troubleshoot complex technical problems spanning multiple systems or services

Preferred qualifications

- Familiarity with incident management tooling and workflows in a large-scale production environment

- Experience with AWS services and cloud infrastructure

- Experience using generative AI or automation to solve operational problems or accelerate workflows

- Track record of authoring post-incident analyses (post-mortems) and driving corrective actions to completion

- Experience building operational dashboards, runbooks, or automation that improved team efficiency

- Knowledge of infrastructure-as-code and deployment tooling, such as CDK, CloudFormation, Terraform, Ansible, or similar

- Experience coordinating across globally distributed teams and time zones

- Comfort operating with ambiguity and incomplete information, and a self-starting approach to identifying what needs doing

Amazon is an equal opportunity employer and does not discriminate on the basis of protected veteran status, disability, or other legally protected status.

Our inclusive culture empowers Amazonians to deliver the best results for our customers. If you have a disability and need a workplace accommodation or adjustment during the application and hiring process, including support for the interview or onboarding process, please visit https://amazon.jobs/content/en/how-we-hire/accommodations for more information. If the country/region you’re applying in isn’t listed, please contact your Recruiting Partner.

The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.

USA, WA, Seattle - 104,500.00 - 160,000.00 USD annually

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Seattle
≈ $84k – $175k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Georgetown
Python
SQL
C#
DevOps
Linux
Windows
Apply
≈ $115k – $212k per year (Estimated) • Remote (United States) • Full-Time • 6+ years exp
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
AI/ML
Prompt Engineering
AI Agents
Frontend
React.js
Apply
$85k – $168k per year • Remote (United States) • 2+ years exp • Bachelor's Degree • Washington
Python
PowerShell
DevOps
Puppet
Chef
Azure
CI/CD
Windows Server
Configuration Management
Hyper-V
Linux
Windows
Unix
DNS
Cybersecurity
Active Directory
Management
OneDrive
Apply
DevOps Engineer 3 days ago
$122k – $152k per year • Remote (United States) • Full-Time • 3+ years exp • PhD • New York
Databases
Azure Cosmos DB
Azure SQL Database
DevOps
Terraform
Azure DevOps
Prometheus
Azure
CI/CD
Kubernetes
Grafana
Azure AKS
FinOps
GitHub
DNS
Cybersecurity
Least Privilege
Apply
$125k – $156k per year • Remote (United States) • Full-Time • United States
DevOps
Terraform
Azure DevOps
Azure
CI/CD
GitOps
Kubernetes
Platform Engineering
Azure AKS
Apply
≈ $55k – $142k per year (Estimated) • In office • Full-Time • Seoul
Python
Verilog
C++
Apply
≈ $55k – $142k per year (Estimated) • In office • Full-Time • Seoul
Python
Verilog
C++
Apply
$135k – $185k per year • Hybrid • 5+ years exp
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Copilot
Claude
Airflow
ChatGPT
dbt
Machine Learning
DevOps
Rest API
GCP
CI/CD
IAM
Analytics
Looker
Apply
In office
Python
SQL
Apply
In office
Python
SQL
Databases
Snowflake
AI/ML
dbt
Apply
$60k – $108k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Ashburn
DevOps
AWS
Linux
Apply
$129k – $175k per year • Equity • In office • TS/SCI • Full-Time • 3+ years exp • Bachelor's Degree • Arlington
Python
Go
Java
Ruby
PowerShell
C#
C++
Databases
Amazon Aurora
AI/ML
Machine Learning
DevOps
AWS
Linux
Management
Agile
Scrum
Apply
$137k – $185k per year • Equity • In office • Full-Time • 6+ years exp • Bachelor's Degree • Seattle
DevOps
AWS
Apply
$111k – $186k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Herndon
DevOps
AWS
Windows
Analytics
Microsoft Excel
Apply
$185k – $250k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Austin
Python
AI/ML
Machine Learning
Chips/EDA
Cadence Allegro
Apply
In office • Full-Time • Bachelor's Degree • Seattle
Apply
≈ $96k – $216k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
Python
SQL
Apply
≈ $164k – $347k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
AI/ML
Edge AI
Apply
≈ $159k – $290k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Seattle
AI/ML
Edge AI
Apply
≈ $155k – $285k per year (Estimated) • In office • Full-Time • Seattle
DevOps
SLI/SLO/SLA
Apply
See all jobs
This is one of many
824,589 more open roles from verified company boards, updated every day.