411,951open jobs
14,442companies
73,805added this week
Browse all
Salary
$150k – $250k per year
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Runloop provides secure code sandboxes and evaluation infrastructure for AI coding agents. Teams launch agents against real repositories, score their output on benchmarks and promote what works into production. The company sells the runtime layer that agent products would otherwise build themselves.

About Runloop

Runloop.ai is pioneering the next generation of infrastructure and orchestration to power the Agentic Web/age of AI Agents. Our platform empowers developers to deploy agents that write code, browse the web, and use computers the way a human would. We're a small team of former Google and Stripe engineers, including the founding team of Google Wallet, dedicated to solving the complex challenges of productionizing AI for software engineering at scale.

The Role

We're looking for a skilled and passionate Site Reliability Engineer to join our team. As a SRE, you'll be responsible for the reliability, observability, performance, and security of our core platform. You'll work closely with our engineering team to develop and maintain the systems that power our code sandboxes, ensuring a seamless and stable experience for our customers. This is a critical role that blends a deep understanding of distributed systems with a software engineering mindset.

Responsibilities

  • Design and maintain our production infrastructure on cloud platforms like AWS, GCP, Azure, and emergent Neo-Clouds

  • Monitor and respond to system alerts and incidents using Grafana and Prometheus, ensuring high availability and a secure environment for our users

  • Collaborate with developers to ensure new features and services are designed with scalability and reliability in mind

  • Troubleshoot and resolve complex issues related to our infrastructure, networking, and the sandbox environment

  • Participate in an on-call rotation to support our production systems

  • Define and track SLIs/SLOs, manage error budgets, and proactively monitor distributed systems with logging and tracing

  • Automate deployments, scaling, provisioning, and recovery tasks to reduce toil and build self-healing systems

  • Lead incident response, conduct root-cause analysis, and facilitate blameless post-mortems to drive continual improvement

  • Collaborate cross-functionally with product, engineering, and developer relations to ensure reliable releases and an outstanding developer experience

  • Plan for capacity growth, forecast system usage, and contribute to safe release and change management processes

Qualifications

  • Strong computer science fundamentals, backed by a degree from a top-tier CS/EE program, or equivalent experience

  • 5+ years of experience in software engineering, with at least 3 years focused explicitly on site reliability, DevOps, or infrastructure operations

  • Strong programming skills in languages like Python or Go

  • Deep expertise in containerization technologies such as Docker and Kubernetes

  • Experience with cloud infrastructure and tools like Terraform and/or Pulumi

  • Familiarity with monitoring and alerting tools like Prometheus, Grafana, or Datadog

  • A solid understanding of networking, security, and Linux systems administration

  • Experience designing, scaling, and maintaining distributed systems (backend platforms, APIs, or front-end infrastructure)

  • Proficiency in implementing observability frameworks (metrics, logging, tracing) and aligning reliability goals with developer velocity

  • Hands-on experience managing incidents, running on-call operations, and producing actionable post-mortems

  • Ability to mentor engineers and influence reliability practices across teams, especially for front-end infrastructure and performance

Bonus Points

  • Experience with chaos engineering techniques, front-end observability tools (e.g., Sentry, RUM, synthetic monitoring), or building CI/CD pipelines for front-end delivery

Benefits

  • Competitive salary and equity

  • Comprehensive health, dental, and vision insurance for employee and dependents

  • Opportunity to work on cutting-edge technology and make a real impact on the future of software engineering

  • Daily catered lunch for all employees and a fridge full of your favorite snacks and drinks

Location:

  • Onsite 4 days a week in San Francisco; Optional 1 day a week remote

Join Us! If you're excited about shaping the future of AI-driven software engineering and empowering developers to build the next generation of AI powered coding tools, we want to hear from you. Join the Runloop team and be at the forefront of the AI revolution in software development.

Runloop AI is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, disability status, protected veteran status, sexual orientation, gender identity, or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
411,951 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$21k – $52k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Hyderabad
Java
TypeScript
JavaScript
Java
Spring Boot
Frontend
Angular
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$28k – $51k per year (Estimated) • Remote/Hybrid • Full-Time • Bucharest
C#
JavaScript
SQL
TypeScript
C#
.NET
Databases
MS SQL
Frontend
Angular
DevOps
Azure
CI/CD
Docker
Kubernetes
Apply
$74k – $187k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sydney
JavaScript
Kotlin
Objective-C
TypeScript
Frontend
React.js
Mobile
Expo
React Native
DevOps
AWS
Azure
GCP
Apply
GTM intern 1 day ago
$20k per year • In office • Internship • 18+ years exp • Paris
SQL
AI/ML
AI Agents
dbt
DevOps
GitHub
Apply
$190k – $230k per year • Equity 0.2–0.8% • In office • Full-Time • 3+ years exp • San Francisco
Rust
AI/ML
AI Agents
DevOps
Kubernetes
Apply
$150k – $250k per year • In office • Full-Time • 4+ years exp • San Francisco
Go
Java
Rust
AI/ML
AI Agents
DevOps
AWS
CI/CD
Configuration Management
Docker
GCP
Grafana
Kubernetes
Prometheus
Pulumi
Terraform
Management
Stripe
Apply
$150k – $250k per year • In office • Full-Time • 4+ years exp • San Francisco
Java
Python
TypeScript
AI/ML
AI Agents
DevOps
AWS
Azure
Docker
GCP
Management
Stripe
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
411,951 more open roles from verified company boards, updated every day.