399,808open jobs
13,937companies
77,696added this week
Browse all
Salary
$300k – $400k per year
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Senior Platform & Reliability Engineer

About OpenArt

OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We’re building the next generation of creative tools powered by cutting-edge AI, enabling anyone to create videos, visuals, characters, and stories with unprecedented speed and imagination.

We believe the future of creativity is AI-native, and we're shaping that future.

Why Join OpenArt

  • Small team, massive surface area, senior engineers own real systems, notslices.

  • Ship at real scale, your work goes to millions of users, fast.

  • Founder-led engineering culture, both founders are technical and deeplyinvolved in product and architecture.

  • AI-native product, you’ll design how cutting-edge AI models are exposed asreal user experiences.

  • High ownership, low process, we value judgment, clarity, and speed overbureaucracy.

  • Senior Platform & Reliability Engineer 1

  • 7-10X growth in revenue for the past 2 years. Now you’ll play a critical role inhelping the company scale to the next stage.

About the Role

We’re looking for a Senior Platform & Reliability Engineer to help design, scale, and improve the reliability of our infrastructure, from architectural decisions to hands-on implementation, observability, and cost optimization.

This is not a traditional ops or DevOps role. You’ll work across cloud infrastructure, distributed systems, backend services, and developer tooling, making pragmatic decisions that balance product velocity, system reliability, and cost efficiency-in a fast-moving, AI-native environment.

You’ll partner closely with product engineers to evolve the platform that powers OpenArt, contributing to key decisions around infrastructure architecture, improving multi-provider AI reliability, and helping us scale systems to millions of users-while raising the overall engineering bar.

What You’ll Do

  • Define and operationalize SLOs/SLIs across critical user journeys (generation, editing, payments/credits, uploads), and use them to guide prioritization and tradeoffs.

  • Participate in an on-call rotation and improve incident response (alert quality, run books, escalation paths), including leading blameless postmortems and driving follow-through on action items.

  • Improve system resilience at external boundaries (AI providers, storage, etc.),including timeouts, retries, circuit breakers, and fallback strategies. Build and maintain end-to-end observability (logs, metrics, traces, dashboards) so engineers can quickly understand “what broke” and “why.”

  • Strengthen deploy safety through CI/CD improvements, automated rollbacks, canary releases, and feature flag patterns.

  • Contribute to the evolution of our infrastructure architecture, helping evaluate when to extend serverless patterns vs. adopt containerized or more managed approaches as we scale.

  • Improve cost visibility and efficiency, including per-request cost attribution, caching strategies, and capacity planning.

  • Act as a strong technical contributor, helping improve engineering practices, tooling, and system design decisions across the team.

What We’re Looking For

Core Requirements

  • 5+ years building and operating production systems where reliability and scaling are important.

  • Strong software engineering skills - you can build and ship production code, not just configure infrastructure.

  • Experience with cloud-native systems (AWS or GCP), including serverless/event-driven architectures and at least one container-based approach (e.g., ECS/Fargate, Cloud Run, Kubernetes).

  • Solid understanding of observability and reliability practices: metrics, alerting, tracing, and incident response.

  • Experience designing resilient systems with external dependencies (timeouts, retries/backoff, idempotency, circuit breakers).

  • Ability to communicate technical tradeoffs clearly to engineers across different domains.

  • Comfortable operating in ambiguous, fast-moving environments and taking ownership of problems.

    Nice to Have

  • Experience building internal platform abstractions (e.g., job orchestration, APIlayers, workflow systems) that improve team velocity.

  • Track record of improving reliability metrics (e.g., MTTR, SLO attainment, latency) or reducing infrastructure cost.

  • Experience working in a startup or high-growth environment, with broad ownership across systems.

Tech Stack You’ll Work With

GCP, Cloud Run, Modal, Upstash, Sentry, Amplitude, Firebase, Redis, React /Next.js, Node.js, TypeScript, Python, etc.

Compensation

  • Competitive base salary and bonus program

  • Equity - meaningful ownership in what you build

  • High autonomy, high growth environment

Work Setup

  • Bay Area preferred (hybrid allowed)

  • Visa sponsorship available

  • We’ll consider remote

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
399,808 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$120k – $150k per year • Equity 1–3% • In office • Full-Time • 1+ year exp • San Francisco
JavaScript
Node JS
Python
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Apply
SRE Engineer 2 hours ago
Remote/Hybrid • 4+ years exp
DevOps
Ansible
AWS
CI/CD
Datadog
Docker
Grafana
Kubernetes
OpenTelemetry
Prometheus
SLI/SLO/SLA
Terraform
Apply
$45k – $180k per year • Equity 0.5–1% • In office • Full-Time • San Francisco
AI/ML
Edge AI
Apply
In office • 4+ years exp • Master's Degree
Python
AI/ML
Image Segmentation
PyTorch
TensorFlow
DevOps
AWS
Docker
GCP
Apply
Mobile Engineer 1 hour ago
$20k – $66k per year (Estimated) • In office • 7+ years exp • Bengaluru
JavaScript
TypeScript
Dart
Frontend
React.js
Mobile
Flutter
React Native
DevOps
CI/CD
Apply
Remote/Hybrid • Full-Time • 3+ years exp • PhD • New York
Apply
$160k – $200k per year • Remote/Hybrid • Full-Time • San Carlos
Apply
$200k – $250k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Carlos
Marketing
Amplitude
Google Ads
HubSpot
LinkedIn
LinkedIn Ads
Meta Ads
Reddit
Apply
SEO & AEO Lead 9 days ago
$200k – $250k per year • In office • Full-Time • 5+ years exp • San Carlos
SQL
Databases
BigQuery
Google BigQuery
AI/ML
Claude
Model Context Protocol
Analytics
Metabase
Marketing
Ahrefs
Amplitude
GA4
Reddit
YouTube
Apply
$130k – $150k per year • In office • Full-Time • San Carlos
Marketing
YouTube
Apply
Founding AE 1 day ago
$130k – $180k per year • Equity 0.2–1% • Remote • Full-Time • 3+ years exp • San Francisco
DevOps
GitHub
Apply
$72k – $120k per year • In office • Internship • San Francisco
Apply
AI/SWE Intern 1 day ago
$36k – $120k per year • In office • Internship • San Francisco
JavaScript
Python
TypeScript
AI/ML
Computer Vision
Frontend
React.js
Apply
$120k – $180k per year • Equity 0.1–0.4% • In office • Full-Time • San Francisco
Python
JavaScript
Frontend
React.js
Apply
$120k – $180k per year • Equity 0.1–0.4% • In office • Full-Time • San Francisco
AI/ML
Computer Vision
LLM
Multimodal AI
Apply
See all jobs
This is one of many
399,808 more open roles from verified company boards, updated every day.