423,509open jobs
14,335companies
62,264added this week
Browse all
Salary
$136k – $272k per year (Estimated)
Location
Remote (United States, Germany, Japan, France, Belgium, Italy, Netherlands, Spain)
Seniority
Staff
Overview
Company
Impact
Profile match
Yuno is an educational technology company based in San Francisco, California, and was founded in 2023. The organization develops an artificial intelligence-driven tutoring platform that provides personalized learning experiences and study assistance for students and professionals. It operates as a digital platform focusing on the integration of large language models to enhance academic performance and knowledge retention.

US · Europe · APAC · Remote · Full Time · Individual Contributor · +7 Years of Experience

Who We Are

At Yuno, we are building the payment infrastructure that allows all companies to participate in the global market. Founded by seasoned experts from the payments and tech industries, our technology provides access to leading payment capabilities, enabling companies to engage customers confidently and maintain global operations through seamless integrations.

We empower high-performing teams at brands like InDrive, McDonald’s, Rappi, and Viva Aerobus to integrate over 1,000 payment methods via a single API. By leveraging advanced AI and the latest technologies, we orchestrate smart routing and fraud prevention across 80+ countries.

About The Role

We are hiring a Staff Incident Manager to own the major incident lifecycle end to end. You are the person who takes control when something is broken in production, brings the right people together, drives the response to resolution, and makes sure we are measurably better after every incident than we were before it.

This is not a passive coordination role. You set the standard for how Yuno detects, responds to, communicates about, and learns from incidents. You will spend as much time fixing the process as you do running the response.

On-call responsibility is core to this role. Yuno's engineering teams run a “You Build It, You Run It” model - pods own their own on-call. The Incident Commander function exists to coordinate cross-domain and Sev-1 events where multiple pods are involved.

Your Contribution Will Be

  • Incident command - act as incident commander on major and critical incidents, owning coordination, decision-making cadence, and escalation from detection through resolution; treat merchant transaction impact, PSP and acquirer dependency failures, settlement and reconciliation knock-on effects, and PCI-DSS scope as first-class concerns in every response.

  • Reliability metrics - drive down time to detect, time to engage, and time to recover; own MTTR as a headline metric and the operational discipline behind Yuno's 99.99% uptime target.

  • On-call program - run rotations, escalation policies, paging hygiene, and alert quality; reduce noise so on-call engineers trust their pages.

  • Incident communications - own internal stakeholder alignment during an event and drive clear, accurate merchant-facing updates - including the status page - in coordination with Support and account teams.

  • Postmortems - run blameless postmortems, hold the room to a no-blame standard, and make sure action items are concrete, owned, and tracked to closure.

  • Reliability roadmap - translate recurring incident patterns into reliability work, partnering with engineering teams to turn postmortem findings into roadmap items, not orphaned tickets.

  • Operating model - define and maintain incident severity levels, response runbooks, and the operating model for declaring and managing incidents.

  • Reporting - report on incident trends, reliability posture, and SLA and SLO performance to engineering leadership.

Skills You Need

Minimum Qualifications

  • Proven experience running major incident response in a production environment that real customers depend on, ideally as an incident commander or in a dedicated incident management function.

  • Strong working knowledge of modern distributed systems and cloud-native operations - comfortable holding your own on a bridge call with senior engineers during an outage, following the technical thread, and keeping the response moving without needing every detail spoon-fed.

  • Experience defining or maturing an incident management practice from the ground up: severity frameworks, operating models, runbooks, and on-call programs built to last, not just inherited.

  • Fluency with observability and incident tooling - metrics, logging, tracing, alerting, and paging platforms. Hands-on experience with Datadog is strongly preferred; familiarity with a paging platform (PagerDuty, OpsGenie, or similar) and a status page tool is expected.

  • A track record of running postmortems that change behavior, and of closing the loop between incidents and engineering work.

  • Excellent written and verbal communication - able to write a clear merchant-facing status update and a precise internal escalation under pressure.

  • Calm, decisive judgment during high-pressure events - you hold structure when others are reacting.

  • Comfort being on-call as a regular part of the role, including for critical incidents outside business hours.

  • English - advanced proficiency required.

Preferred Qualifications

  • Familiarity with PCI-DSS and the operational obligations that come with handling payment flows at scale.

  • Spanish - business-level proficiency preferred; Yuno's engineering organization spans LatAm, and on-call bridges and postmortem discussions often run in Spanish.

Nice to Have

  • Prior experience in payments, fintech, or another high-availability, regulated domain.

  • Exposure to SRE practices, error budgets, and SLO-driven prioritization.

  • Experience managing incidents across multi-timezone, remote-first engineering organizations.

What Success Looks Like

  • First 90 days - you have mapped how incidents currently get declared, run, and reviewed, identified the biggest gaps, and started closing them. Severity definitions and the incident operating model are clear and adopted.

  • Six months - major incidents run to a consistent standard. Postmortem action items are tracked and closed. On-call noise is down and trust in paging is up. Leadership has a reliable view of reliability trends.

  • One year - the incident practice is self-sustaining. Runbooks are owned by engineering pods, not just by you. MTTR and repeat-incident rates are trending measurably down, and reliability work sourced from incident data is landing on engineering roadmaps and shipping.

What We Offer at Yuno

  • Competitive Compensation.
  • Remote Work - You can work from everywhere!
  • Home Office Bonus - A one-time allowance to help you create your ideal home office.
  • Work Equipment.
  • Stock Options.
  • Health Plan wherever you are.
  • Flexible Days Off.
  • Language, Professional, and Personal Growth courses.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
423,509 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$43k – $105k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Prague
Python
Python
Flask
FastAPI
DevOps
Terraform
GitHub Actions
OpenTelemetry
CloudFormation
Datadog
Prometheus
GitLab CI
CI/CD
Git
AWS
Docker
Kubernetes
Grafana
Amazon EKS
AWS Lambda
Amazon EC2
AIOps
GitHub
GitLab
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Management
Confluence
Apply
$20k – $58k per year (Estimated) • Equity • Remote • 2+ years exp • Bachelor's Degree
AI/ML
AI Agents
Apply
$108k – $243k per year (Estimated) • Equity • Remote • 5+ years exp
AI/ML
AI Agents
Apply
$82k – $184k per year (Estimated) • Equity • Remote • 8+ years exp
Python
Go
Kotlin
Databases
PostgreSQL
Redis
Snowflake
Weaviate
Databricks
Pinecone
FAISS
Apache Kafka
Google BigQuery
BigQuery
Kafka
AI/ML
Airflow
Model Context Protocol
MLFlow
Prefect
Fine-tuning
RLHF
Prompt Engineering
Function Calling
AI Agents
Langfuse
LangSmith
LLM
RAG
Braintrust
Feature Store
Text-to-Speech
LLM Guardrails
Tool Use
DevOps
Terraform
GitHub Actions
OpenTelemetry
Datadog
CI/CD
ArgoCD
Git
AWS
Docker
Kubernetes
Vector
AIOps
GitHub
Management
WhatsApp
Apply
$81k – $172k per year (Estimated) • Equity • Remote • 8+ years exp
Go
Java
Kotlin
Java
Spring Boot
Databases
PostgreSQL
Redis
Apache Kafka
Kafka
AI/ML
AI Agents
DevOps
gRPC
Terraform
GitHub Actions
OpenTelemetry
Datadog
CI/CD
ArgoCD
Git
AWS
Docker
Kubernetes
Amazon EKS
GitHub
Amazon S3
Amazon ECS
Cybersecurity
SonarQube
Apply
$81k – $172k per year (Estimated) • Equity • Remote • 8+ years exp
Go
Java
Kotlin
Java
Spring Boot
Databases
PostgreSQL
Redis
Apache Kafka
Kafka
AI/ML
AI Agents
Tokenization
DevOps
gRPC
Terraform
GitHub Actions
OpenTelemetry
Datadog
CI/CD
ArgoCD
Git
AWS
Docker
Kubernetes
Amazon EKS
GitHub
Amazon S3
Amazon ECS
Cybersecurity
SonarQube
PCI DSS
Apply
See all jobs
This is one of many
423,509 more open roles from verified company boards, updated every day.