368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$43k – $108k per year (Estimated)
Location
Remote (Spain)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Yuno is an educational technology company based in San Francisco, California, and was founded in 2023. The organization develops an artificial intelligence-driven tutoring platform that provides personalized learning experiences and study assistance for students and professionals. It operates as a digital platform focusing on the integration of large language models to enhance academic performance and knowledge retention.

Remote · Full Time · Individual Contributor · +7 Years of Experience

Site Reliability Engineer

Who We Are

Yuno is the AI-native operating system of global commerce, powering the financial infrastructure of enterprise merchants, banks, and wallets. Through a single API, Yuno connects them to pay-ins, payouts, fraud prevention, KYC/KYB, and stablecoins globally, so they can operate everywhere. Agnostic by design and connected to 1,000+ payment methods and 460+ integrations in 190+ countries, Yuno optimizes acceptance rates, reduces costs, and strengthens security through specialized AI agents that learn from every transaction. Global brands including McDonald's, NetEase Games, GoFundMe, and Rappi run their payments on Yuno.

About The Role

Yuno is looking for a Staff Site Reliability Engineer to set the technical direction for reliability across our infrastructure - starting with the platform that provisions, deploys, and manages AI agents at scale on AWS, the system powering payments across 190+ countries. The platform is in production and growing, and we need the most senior reliability voice in the room to evolve the architecture and make sure it stays reliable, observable, and ready to scale.

This is not a "maintain what exists" role, and it's not a single-system role. You'll own the reliability strategy - driving architectural decisions, designing event-driven communication, defining how we measure and defend reliability, and setting the standards other engineering teams build on.

How AI Shows Up in This Role

  • The platform you own is Yuno's AI agent infrastructure - provisioning and deploying AI agents at scale, plus the agents that route payments and prevent fraud. Keeping the AI-native layer reliable is the core of the role

  • AI is our default execution layer: you're encouraged to use AI-assisted tooling across automation, runbooks, incident analysis, and root-cause investigations, and to help define how the wider engineering org adopts it. We care how you use it, not whether you do

Your Contribution Will Be

  • Reliability strategy and standards - define the SLO culture, error-budget policy, and incident practices that scale across engineering teams, turning reliability from firefighting into a measurable, org-wide discipline

  • Platform architecture and evolution - drive architectural decisions as the platform matures; the deciding voice on choosing technologies, designing systems, and when to evolve the infrastructure

  • Messaging and event-driven architecture - design and own the messaging layer for inter-service communication, replacing synchronous patterns with durable, reliable async messaging

  • Infrastructure and deployment - own the cloud infrastructure, automate provisioning with IaC, and ensure the platform scales reliably as transaction volume grows

  • Observability - build the monitoring, tracing, and alerting that keeps the platform healthy; when something breaks at 3am, your dashboards and alerts should explain why before anyone has to dig

  • Incident leadership and mentorship - act as the senior escalation point for the hardest production problems, run blameless postmortems and root-cause analyses that turn into permanent fixes, and raise the reliability bar by mentoring senior and mid-level engineers

  • Chaos engineering mindset - continuous fault injection and resilience experiments that surface weaknesses before they turn into incidents, plus identifying and proposing resilience patterns to prevent those failures from reaching production.

What Success Looks Like

Within your first 6-12 months, you've set the reliability strategy for the platform, driven at least one major architectural evolution (event-driven messaging, streaming reliability, or observability), and engineering teams have adopted the SLO and error-budget framework you defined. You're the person Yuno trusts with the hardest reliability calls.

Skills You Need

Minimum Qualifications

  • Event-driven architecture and messaging systems - you've designed and owned systems around message queues (Kafka, NATS, RabbitMQ) and understand at-least-once delivery, consumer groups, dead letters, and backpressure; you've migrated a system from synchronous to async

  • Deep AWS - EC2, VPC, IAM, S3, and RDS - with strong networking fundamentals, since inter-service communication runs over the internal VPC

  • Infrastructure as Code - Terraform or Pulumi, reviewed in PRs rather than clicked in consoles

  • Kubernetes and Docker in production - container lifecycle, resource limits, health checks, and orchestration at scale

  • Observability and SLOs - Datadog fluency or equivalent (dashboards, monitors, APM, distributed tracing), and a track record defining and operating SLOs, SLIs, and error budgets across services

  • Chaos engineering and resilience testing - hands-on experience with fault injection, game days, or chaos experiments (Gremlin, Chaos Mesh, AWS FIS, or similar) to harden production systems

  • Distributed systems debugging - you've diagnosed async flows and cascading failures in production and can explain what broke and how you fixed it; comfortable coding for automation and tooling (Go, Python, or similar)

  • Databases - solid SQL (PostgreSQL) and NoSQL (MongoDB, Redis): when to use each, indexing, replication, and performance tuning

  • Proven technical leadership - you've set reliability standards, influenced architecture across teams, and mentored engineers, not just owned your own scope

  • English - advanced proficiency, written and spoken

Preferred Qualifications

  • AI / MLOps infrastructure - running AI workloads in production (model serving, LLM inference, GPU/resource management, and agent evaluation/observability tools like LangFuse, LangSmith, Braintrust, or MLflow)

  • Multi-tenant container platforms - running customer or user workloads in containers (Replit, Railway, Fly.io, or internal PaaS)

  • Data pipelines and orchestration - Airflow, Prefect, or similar; data warehouses like Databricks, Snowflake, or BigQuery a plus

  • Incident management and on-call tooling - PagerDuty, Opsgenie, or incident.io

  • Experience in the payments industry

Nice to Have

  • ECS experience

  • s6-overlay for container process supervision

  • Experience with AI agent framework ecosystems

  • Spanish proficiency

What We Offer at Yuno

  • Competitive Compensation

  • Remote Work - you can work from everywhere

  • Home Office Bonus - a one-time allowance to set up your ideal home office

  • Work Equipment

  • Stock Options

  • Health Plan wherever you are

  • Flexible Days Off

  • Language, Professional, and Personal Growth courses

Disclaimer

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed or wish to exercise your data protection rights, please contact us at [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Spain
$142k – $215k per year • In office • Full-Time • 12+ years exp • Princeton
JavaScript
TypeScript
Java
Java
Gradle
Hibernate
Maven
Spring Boot
Databases
Databricks
Snowflake
AI/ML
AI Agents
Claude
Claude Code
Frontend
Angular
React.js
Vue.js
DevOps
Amazon CloudWatch
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
API Gateway
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
IAM
Platform Engineering
Rest API
Kubernetes
Apply
$135k – $190k per year • In office • Full-Time • 12+ years exp • New York • Princeton
Databases
Apache Kafka
AI/ML
Hadoop
PyTorch
TensorFlow
DevOps
AWS
Azure
Azure DevOps
CI/CD
GCP
GitLab
GitLab CI
Jenkins
QA
Appium
Cypress
Playwright
Postman
Rest-Assured
Selenium
Apply
$99k – $186k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Lincoln • Alpharetta
AI/ML
Claude
AI Agents
OpenAI
DevOps
AWS
Azure
Docker
Git
Kubernetes
Rest API
Management
Confluence
QA
Postman
Swagger
Apply
In office • Full-Time • Gurgaon
DevOps
AWS
GCP
SLI/SLO/SLA
Apply
$305k per year • In office • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
AI Agents
Anthropic
Claude
LLM
Multimodal AI
Apply
$43k – $108k per year (Estimated) • Equity • Remote • 7+ years exp
Python
SQL
Databases
Apache Kafka
Databricks
Google BigQuery
NATS
PostgreSQL
RabbitMQ
Redis
Snowflake
AI/ML
AI Agents
Braintrust
Langfuse
LangSmith
LLM
MLFlow
Prefect
Replit
DevOps
Amazon EC2
AWS
Chaos Engineering
Datadog
Docker
Fly.io
Incident Management
Kubernetes
Opsgenie
PagerDuty
Pulumi
SLI/SLO/SLA
Terraform
Amazon ECS
Amazon S3
IAM
Apply
$74k – $154k per year (Estimated) • Equity • Remote • 8+ years exp
Go
Java
Kotlin
Python
Java
Spring Boot
Databases
Apache Kafka
PostgreSQL
Redis
AI/ML
AI Agents
DevOps
Amazon EKS
ArgoCD
AWS
CI/CD
Datadog
Docker
Git
GitHub Actions
Kubernetes
OpenTelemetry
Terraform
Amazon ECS
Amazon S3
GitHub
Cybersecurity
SonarQube
Apply
Engineering Manager 28 days ago
$68k – $152k per year (Estimated) • Equity • Remote • 4+ years exp
Go
Java
Kotlin
SQL
AI/ML
AI Agents
Claude
Claude Code
Apply
$56k – $127k per year (Estimated) • Remote • Full-Time • 4+ years exp • Spain
Java
Kotlin
Java
Spring Boot
Databases
Apache Kafka
PostgreSQL
Redis
DevOps
ArgoCD
AWS
CI/CD
Datadog
Docker
GCP
Git
GitHub Actions
Kubernetes
OpenTelemetry
Terraform
GitHub
Cybersecurity
PCI DSS
Apply
$101k – $182k per year (Estimated) • Equity • Remote • 5+ years exp
Go
JavaScript
Node JS
Python
Databases
MySQL
Oracle
PostgreSQL
DevOps
Incident Management
Nginx
QA
Postman
SoapUI
Swagger
Apply
Process Engineer 6 hours ago
$30k – $78k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Spain
Apply
$40k – $91k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Madrid
Design
Figma
Apply
Product Engineer-2 5 days ago
Remote/Hybrid • Full-Time • Spain
Apply
$30k – $79k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Getafe • Munich
DevOps
Kubernetes
OpenShift
Red Hat
Amazon S3
Apply
$28k – $69k per year (Estimated) • Remote • Full-Time • 3+ years exp • Spain
Apex
Apex
Visualforce
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.