561,906open jobs
22,097companies
76,904added this week
Browse all
Salary
$210k – $240k per year
Location
In office (San Francisco)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About the Role

We’re looking for an experienced Site Reliability Engineer (SRE) to help us scale our platform with reliability, observability, and operational excellence at the core. You’ll partner with engineers and data scientists to build, automate, and maintain the infrastructure that powers our core platform-including data pipelines, ML workloads, and real-time analytics systems.

This is a hands-on, high-impact role with visibility across the stack and the opportunity to shape the future of our infrastructure and operations.

Key Responsibilities

  • Design, build, and maintain scalable infrastructure to support real-time analytics and machine learning workloads

  • Improve system reliability and performance through automation, observability, and proactive capacity planning

  • Own and evolve CI/CD pipelines, deployment automation, rollback mechanisms, and config management

  • Implement and maintain monitoring, alerting, and incident response processes (SLOs, runbooks, on-call rotations)

  • Collaborate across engineering and data science teams to drive a culture of performance and reliability

  • Ensure security, compliance, and operational readiness across our cloud infrastructure

  • Drive post-incident analysis and continuous improvement initiatives

What Will Help You Succeed

  • 8+ years of experience in SRE, DevOps, or infrastructure engineering roles

  • 5+ years of experience with datacenter operations and/or system and network administration

  • Experience with containerization (Docker), and orchestration (Kubernetes)

  • Strong knowledge of Linux systems, networking, and systems performance tuning, including strong-to-expert facility with SSH, terminal, shell and the Linux command line.

  • Solid understanding of infrastructure-as-code (e.g., Terraform, Ansible)

  • Good programming skills and ability to apply sound coding principles to IaC and scripting code with languages such as Terraform, Ansible, Bash (shell scripting), and/or Python.

  • Experience with monitoring and observability stacks (e.g., Prometheus, Grafana, Datadog, ELK, OpenTelemetry) Proficiency with CI/CD tools and pipelines (e.g., GitHub Actions, ArgoCD, etc.)

  • Ability to debug complex systems and automate solutions in scripting languages

  • Excellent communication skills and the ability to work cross-functionally

Nice-to-Have

  • Experience with cloud and managed services (e.g. AWS)

  • Experience supporting data-intensive platforms (Spark, Airflow, Kafka, etc.)

  • Familiarity with security practices for cloud-native applications and infrastructure

  • Experience in high-compliance or SOC-2 environments

What You’ll Get

  • Ownership of mission-critical infrastructure in a company solving real-world enterprise problems

  • A front-row seat to a high-performance engineering culture

  • The ability to influence how our platform scales-from deployment to incident management

  • An environment that values curiosity, accountability, and impact

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
561,906 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$18k – $43k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pune
Python
AI/ML
Copilot
AI Agents
OpenAI
DevOps
Rest API
Terraform
Ansible
Azure
AWS
AIOps
Management
ServiceNow
Apply
$19k – $45k per year (Estimated) • In office • Full-Time • 5+ years exp • Indore
Python
DevOps
Rest API
Terraform
Ansible
CI/CD
Git
Cybersecurity
Zscaler
Apply
$17k – $42k per year (Estimated) • Remote • Full-Time • Moscow
Python
Bash
Databases
OpenSearch
AI/ML
Model Context Protocol
vLLM
Ollama
LLM
DevOps
Terraform
Docker Compose
Helm
Prometheus
Yandex Cloud
GitLab CI
CI/CD
GitOps
ArgoCD
Docker
Kubernetes
Grafana
Harbor
GitLab
Cybersecurity
SBOM
Apply
$26k – $61k per year (Estimated) • Remote/Hybrid • 7+ years exp • Bachelor's Degree • Pune
Python
SQL
Scala
Databases
Snowflake
Apache Kafka
AI/ML
Cursor
Spark
Claude Code
OpenAI Codex
DevOps
GitHub Actions
CI/CD
Jenkins
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon CloudWatch
Amazon Kinesis
Apply
$68k – $196k per year (Estimated) • Remote/Hybrid • Full-Time • Sydney
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
DevOps
Rest API
Terraform
Azure DevOps
Azure
CI/CD
Git
Bicep
Apply
$140k – $186k per year • In office • Full-Time • San Francisco
Python
Python
Alembic
Cybersecurity
Okta
Least Privilege
Management
Google Workspace
Apply
Personal Trainer 22 days ago
$95k – $115k per year • In office • Full-Time • San Francisco
Apply
Technical Recruiter 2 months ago
$125k – $160k per year • In office • Full-Time • San Francisco
C++
AI/ML
CUDA Toolkit
CUDA
Apply
Technical Sourcer 2 months ago
$110k – $135k per year • In office • Full-Time • PhD • San Francisco
C++
AI/ML
CUDA Toolkit
CUDA
Semantic Search
DevOps
GitHub
HPC
Marketing
LinkedIn
Apply
$210k – $240k per year • In office • Full-Time • 8+ years exp • San Francisco
Python
Bash
Python
Alembic
Databases
Apache Kafka
AI/ML
Spark
InfiniBand
DevOps
Terraform
Ansible
OpenTelemetry
Datadog
Prometheus
Kubernetes
Grafana
Configuration Management
Incident Management
Cybersecurity
SOC 2
Apply
$106k – $130k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Chicago • Scottsdale
DevOps
GCP
Azure
CI/CD
AWS
Harbor
Platform Engineering
Incident Management
Apply
$186k – $345k per year (Estimated) • In office • Full-Time • 12+ years exp • Master's Degree • San Francisco
DevOps
Incident Management
Apply
$160k – $236k per year • Equity • Remote/Hybrid • Full-Time • 2+ years exp • Chicago • Austin • San Francisco
Apply
$40k – $100k per year • In office • Full-Time • 1+ year exp • San Francisco
AI/ML
Claude
AI Agents
DevOps
Azure
Analytics
A/B Testing
Design
Adobe Photoshop
Figma
Marketing
Google Ads
Apply
$74k – $219k per year • Remote/Hybrid • Full-Time • 3+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Management
Agile
Apply
See all jobs
This is one of many
561,906 more open roles from verified company boards, updated every day.