612,693open jobs
33,759companies
85,932added this week
Browse all
Salary
$103k – $224k per year (Estimated)
Location
In office (Santa Clara)
Seniority
Senior · 6+ years exp
Overview
Company
Impact
Profile match

Forward was founded in 2013 by four Stanford Ph.D.s, building the industry's first network digital twin: a mathematically accurate model of the production network. It's the foundation for autonomous networking, giving engineers and AI agents the ability to know the impact of every change before it touches production. That founding instinct still defines how we work. We're accurate and evidence-driven, relentless about clarity, and we'd rather be certain than comfortable, building a groundbreaking platform that transforms how teams run and secure networks across every major cloud and vendor environment.

Global leaders like Goldman Sachs, PayPal, S&P Global, IBM, and Dell trust Forward, alongside fast-growing enterprises and government agencies, realizing an average of $14.2 million in annual benefits, according to IDC. Backed by top-tier investors, including A. Capital, Andreessen Horowitz, Goldman Sachs, MSD Partners, Omega Venture Partners, Section 32, and Threshold Ventures, and headquartered in Santa Clara, we're most proud of our team: curious people who'd rather build what doesn't exist than accept how things have always been done.

Forward is looking for a Site Reliability Engineer

About the Role This is not a "keep the lights on" SRE role. As our first or early SRE hire you will be building the reliability engineering function at Forward - defining how we think about availability, observability, incident response, and operational excellence across a complex, distributed SaaS platform. You will work closely with engineering, infrastructure, and product to ensure our platform meets the reliability bar our enterprise customers demand.

If you thrive in environments where you're handed a problem rather than a playbook this role is for you.

What You'll Own

  • Define and drive SRE practices from the ground up - SLOs, SLIs, error budgets, and the frameworks the engineering org will actually use
  • Drive the reliability and operational excellence of the Forward SaaS platform
  • Build and maintain observability infrastructure - logging, metrics, tracing, and alerting - so the team always knows what's happening before customers do
  • Lead incident response: on-call rotations, runbooks, post-mortems, and the follow-through to make sure the same incident doesn't happen twice
  • Partner with engineering teams to embed reliability thinking into the SDLC - capacity planning, load testing, chaos engineering, and production readiness reviews
  • Help define and build the SRE team as the company scales - this is a foundational hire with a path to leadership

What We're Looking For

  • 6+ years of experience in site reliability engineering, DevOps, or infrastructure engineering in a SaaS or cloud environment
  • Proven experience building or significantly maturing an SRE function - not just operating within one someone else built
  • Strong fundamentals in networking - TCP/IP, DNS, routing, switching, firewalls, and load balancing. Experience with network management or observability platforms is a significant plus
  • Hands-on experience with Kubernetes and container orchestration in production environments
  • Deep proficiency with observability tooling - Prometheus, Grafana, Datadog, Splunk, or similar
  • Strong scripting and automation skills in Python, Bash, or similar
  • Experience with cloud platforms - AWS, GCP, or Azure - including infrastructure as code (Terraform, Ansible, or equivalent)
  • Track record of owning and improving incident response processes including blameless post-mortems and SLO-driven reliability improvements
  • Ability to communicate clearly with both engineering teams and non-technical stakeholders - you can explain an outage to a customer-facing team without jargon and explain an SLO to an executive without losing them

Nice to Have

  • Experience supporting enterprise or federal government customers with high availability requirements
  • Experience in a foundational or early SRE hire capacity at a growth stage company

What This Role Is Not

  • A pure ops or NOC role - you are building and engineering, not just monitoring
  • A siloed function - you will be deeply embedded with product and engineering teams
  • A ticket-taker - you will be proactively identifying and solving reliability problems before they become incidents

Why Forward

  • You'll be building something from scratch at a company with real enterprise traction and world-class investors behind it
  • Our customers include some of the most complex network environments on the planet - the reliability bar is high and the work is genuinely interesting
  • People-centric culture built by Stanford Ph.D.s who care deeply about doing things the right way
  • Competitive compensation, equity, and the opportunity to grow into a leadership role as the SRE function scales

The base pay range for this role is between $230,000 and $250,000. Base pay will depend on your skills, qualifications, experience, and location

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
612,693 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
Python
PowerShell
ABAP
ABAP
SAP BTP
Databases
SAP HANA
DevOps
Ansible
GCP
Azure
AWS
Apply
$100k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
Python
Go
JavaScript
Java
Node JS
Bash
Node JS
Commander.js
DevOps
GCP
Istio
OpenTelemetry
Consul
Datadog
Linkerd
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Chaos Engineering
Service Mesh
Apply
Tech Lead - Java 1 day ago
$26k – $65k per year (Estimated) • In office • 12+ years exp • Bengaluru
JavaScript
Java
SQL
Java
Maven
Spring Boot
Hibernate
Spring MVC
Spring Data JPA
Gradle
Databases
MySQL
PostgreSQL
Oracle
Apache Kafka
DevOps
GCP
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
Mobile QA Engineer 1 day ago
$14k – $30k per year (Estimated) • Remote • 3+ years exp • Saint Petersburg
Python
JavaScript
SQL
AI/ML
Copilot
ChatGPT
Mobile
Maestro
DevOps
CI/CD
QA
Appium
Postman
Apply
$132k – $238k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Fremont
Python
JavaScript
Java
Ruby
Node JS
Databases
MySQL
PostgreSQL
Redis
DevOps
Azure
AWS
Cybersecurity
Okta
theHarvester
Apply
$140k – $311k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • PhD • Santa Clara
JavaScript
Java
AI/ML
AI Agents
Frontend
Svelte
Web3
Layer 2
Robotics
Digital Twin
Apply
Engineering Manager 4 days ago
$239k – $249k per year • In office • 2+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
Platform Engineering
Robotics
Digital Twin
Apply
Senior Accountant 10 days ago
$110k – $130k per year • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
Robotics
Digital Twin
Apply
$78k – $147k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree
Python
AI/ML
AI Agents
Frontend
GraphQL
Robotics
Digital Twin
Apply
$129k – $268k per year (Estimated) • In office • 6+ years exp • PhD • Santa Clara
AI/ML
AI Agents
Robotics
Digital Twin
Apply
$140k – $311k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • PhD • Santa Clara
JavaScript
Java
AI/ML
AI Agents
Frontend
Svelte
Web3
Layer 2
Robotics
Digital Twin
Apply
$145k – $314k per year (Estimated) • Remote/Hybrid • Contractor • Master's Degree • Santa Clara
Python
AI/ML
llama.cpp
LoRA
vLLM
Fine-tuning
Quantization
Multimodal AI
Knowledge Distillation
AI Agents
VLM
SGLang
GGUF
TensorRT
PEFT
TensorRT-LLM
Transformers
PyTorch
LLM
Mixture of Experts
DPO
Post-training
Edge AI
Speculative Decoding
KV Cache
Multi-Agent Systems
Model Distillation
Analytics
ETL/ELT
Management
Freshdesk
Apply
$201k – $352k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
TypeScript
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
AI Agents
LLM
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
Function Calling
AI Agents
RAG
OpenAI Codex
Human-in-the-Loop
LLM Guardrails
Multi-Agent Systems
Tool Use
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 15+ years exp • Bachelor's Degree • Santa Clara
Python
Go
Java
AI/ML
Cursor
Windsurf
Claude Code
AI Agents
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
See all jobs
This is one of many
612,693 more open roles from verified company boards, updated every day.