368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$180k – $220k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Vapi is an artificial intelligence technology company headquartered in San Francisco, California, and founded in 2020. The company provides a developer platform for building, testing, and deploying conversational voice AI agents that integrate speech-to-text, large language models, and text-to-speech technologies. It serves enterprise clients across sectors such as healthcare, finance, and retail, offering scalable infrastructure for automating customer support and sales operations.

Voice AI that resolves, not transfers.

Most phone systems trap callers in menus and scripts. Vapi is the platform for deploying voice agents that know your business and can listen, adapt, and resolve in minutes.

  • The numbers: 1 billion calls. 1 million developers. 10x enterprise ARR growth

  • The customers: Amazon Ring, ServiceTitan, New York Life, Intuit, Kavak, and thousands more, from YC startups to the Fortune 500

  • The news: a $50M Series B led by Peak XV Partners, with Bessemer Venture Partners, Kleiner Perkins, M12 (Microsoft's Venture Fund), Y Combinator, and our earlier backers. Total raised: $72M

Incident and Escalation Manager

Why this role exists

Vapi runs voice AI infrastructure at tens of millions of minutes a month for enterprise customers who route business-critical traffic through us. When something breaks, it breaks in real time, in front of a customer's own customers, and often through a carrier dependency we don't fully control. Today the response is run by whoever is in the Slack thread. That works once. It does not scale, and it costs us renewals, engineering focus, and executive confidence every time it happens.

This is the first dedicated hire for the function, and the job is to build it. You will design the program that runs our incident response, train the people who command incidents, and own the customer relationship through the weeks that follow a serious one. You own the pager and you work part of the rotation yourself, because the fastest way to build a program that works is to run inside it while you build it. You are not the engineer fixing the system. You are the person who builds the system so a trained bench can run it, and who takes a real turn on the pager.

Building the program

This is the core of the role.

  • Define what counts as an incident and what does not, with entrance and exit criteria that protect engineering from noise.

  • Author the severity model with response-time targets per level, so a SEV0 means the same thing Monday morning and Friday night.

  • Design how a live incident runs: the command structure, the update cadence, the single source of truth, the decision rights, and when the call gets escalated past the commander. Write it down so anyone on rotation runs it the same way.

  • Build and train the incident commander rotation. Build a realistic balance of hiring and utilizing existing resources across humans and AI to stand up the rotation.

  • Own the pager and work part of the rotation yourself - personally commanding incidents during your shifts and stepping your share down as the bench matures and proves out. Own how the rotation runs regardless of whose shift it is.

  • Stand up the incident tooling and on-call setup: paging, escalation policies, incident channels, status page, and the runbook library.

  • Build customer communication templates for each severity and channel, pre-approved so they are not written from scratch under pressure.

  • Govern the customer credit process with a clear approval chain, so financial decisions stop happening in ad hoc threads.

  • Stand up the metrics: resolution time, response velocity, escalation volume, RCA SLA adherence, and revenue protected. Report monthly in terms execs use to make resource decisions.

  • Build standing partnerships with engineering, support, the office of the CTO, legal, security, comms, and carrier operations before the next critical situation.

  • Train go-to-market, support, customer success, and engineering on where to bring customer-critical issues and how the function works.

  • Build a feedback loop so incident and escalation data shapes the engineering roadmap instead of dying in postmortems.

After the incident

  • Own the customer-facing RCA. Translate engineering root cause into plain language that tells the truth and holds the relationship. Ship it within the SLA we commit to.

  • Run the blameless post-incident review. Drive action items to named owners with dates, and track them to closure instead of letting them die in a doc.

  • Close the loop with affected customers directly, including the credits or commitments made during the incident.

Escalation management

  • Hold the high-severity customer issues that do not rise to a full incident but threaten a renewal or a relationship.

  • Run the standing executive escalation list. Keep an owner, a next step, and a date on every item.

  • Spot patterns across customers that no single team owns, and force them into engineering or product as prioritized work.

  • Be the single point of contact for a regulated customer or a regulatory inquiry that surfaces weeks after the technical fix.

  • Partner with the account team on at-risk accounts driven by reliability, and build the cross-functional recovery plan.

  • Keep a written handoff and a named deputy so escalations stay covered when you are out.

What you bring

  • 8 to 12 years across incident management, escalation management, technical support escalations, or technical program management, ideally at an infrastructure, telephony, or platform company operating at scale.

  • A track record building an incident and escalation program from zero, or owning a meaningful piece of one through its growth, including standing up an incident commander rotation rather than being the sole responder. Experience hiring a team is a plus.

  • Hands-on familiarity with incident tooling such as PagerDuty, incident.io, Opsgenie, or equivalent, and the Slack and status-page workflows around them.

  • Calm under pressure as a learned discipline. Gravitas to direct a response and the willingness to remove a distraction from a call even when it outranks you.

  • Decision-making with incomplete information, and the judgment to know when to escalate and how to do it without losing time.

  • Clear writing under pressure. You can produce a clean read of a live situation for a senior leader, and a customer RCA that holds a relationship together.

  • Comfort with technical depth. You do not need to write the fix, but you need to follow the conversation, ask the right question, and know when an answer does not add up. Familiarity with telephony, carrier dynamics, or real-time systems is a strong plus.

  • Willingness to work part of the incident commander rotation, including off-hours shifts, especially in the first year.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
In office • Full-Time • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Databases
Oracle
Frontend
Angular
DevOps
Incident Management
SLI/SLO/SLA
Apply
$92k – $199k per year (Estimated) • Remote/Hybrid • Full-Time • Australia
DevOps
SLI/SLO/SLA
Apply
In office • Full-Time • Gurgaon
DevOps
AWS
GCP
SLI/SLO/SLA
Apply
$17k – $21k per year (Estimated) • Remote/Hybrid • Full-Time • Bucharest
DevOps
Azure
GCP
Incident Management
Splunk
Cybersecurity
Google SecOps
MITRE ATT&CK
Apply
$74k – $196k per year (Estimated) • In office • Full-Time • 8+ years exp • Dublin
C#
Java
Node JS
Python
JavaScript
Java
Hibernate
Spring Boot
DevOps
AWS
Azure
GCP
Cybersecurity
CWE
Apply
$200k – $280k per year • Equity • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Node JS
Python
SQL
TypeScript
JavaScript
Python
Django
FastAPI
Flask
Databases
DynamoDB
pgvector
Pinecone
Redis
PostgreSQL
AI/ML
Claude
Claude Code
Cursor
LLM
RAG
Windsurf
OpenAI Codex
Structured Outputs
Text-to-Speech
Vapi
Model Context Protocol
Frontend
Next.js
React.js
Tailwind CSS
Vue.js
DevOps
AWS
Azure
Docker
GCP
GitOps
Kubernetes
Vector
Vercel
WebRTC
WebSockets
Apply
$200k – $325k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
Go
Node JS
TypeScript
JavaScript
Node JS
BullMQ
Databases
Apache Kafka
PostgreSQL
AI/ML
AI Agents
LLM
LLM Guardrails
Text-to-Speech
Vapi
DevOps
Amazon EKS
AWS
Kubernetes
Apply
$180k – $265k per year • Equity • Remote/Hybrid • Full-Time • San Francisco
Node JS
TypeScript
JavaScript
Java
Node JS
BullMQ
Nest.JS
Java
Liquibase
Databases
Apache Kafka
PostgreSQL
AI/ML
LLM
LiveKit
Text-to-Speech
Vapi
DevOps
OpenTelemetry
Management
Slack
Apply
$211k – $428k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • San Francisco
Go
Python
TypeScript
AI/ML
Claude
Claude Code
Cursor
LLM
OpenAI Codex
Vapi
AI Agents
Model Context Protocol
DevOps
CI/CD
Apply
Senior Brand Designer 14 days ago
$200k – $230k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
AI/ML
Vapi
Design
Figma
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.