619,872open jobs
30,169companies
85,569added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Overview

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware, a position that took years to build.

About the Role

We're looking for a Site Reliability Engineer to help make vLLM-powered inference systems reliable, observable, and operationally simple at production scale. This role is for someone who thinks about failure before launch, designs systems that are easier to operate, and knows how to turn incidents into durable improvements rather than one-off fixes.

You'll work across engineering and infrastructure to define SLOs, improve monitoring and alerting, strengthen incident response, drive post-mortems, and reduce operational risk before it reaches users. Your work will directly impact the reliability, availability, and production readiness of the systems powering AI inference at scale.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems, infrastructure, or similar.

  • Strong experience operating production systems with meaningful traffic, user impact, or infrastructure criticality.

  • Deep understanding of SLOs, SLIs, error budgets, alerting, incident response, and post-mortem processes.

  • Experience live-fighting major production incidents, including mitigation, root cause analysis, escalation, and follow-through on prevention work.

  • Strong Linux, networking, systems debugging, observability, and distributed systems fundamentals.

  • Ability to design operationally simple systems and identify likely failure modes before launch.

  • Strong programming or scripting ability in Python, Go, Bash, or similar for automation, tooling, and reliability improvements.

Preferred qualifications:

  • Experience supporting ML infrastructure, inference systems, GPU workloads, Kubernetes-based platforms, or high-scale backend services.

  • Experience building or improving observability systems using metrics, logs, traces, dashboards, alerts, and runbooks.

  • Experience with Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, CI/CD systems, or production deployment platforms.

  • Experience driving incident review culture, post-mortem processes, reliability reviews, and prevention-oriented engineering work.

  • Ability to partner with engineering teams to improve service design, release safety, capacity planning, and operational readiness.

Bonus points if you have:

  • Owned reliability for high-throughput, latency-sensitive, or mission-critical production systems.

  • Supported AI inference, model serving, GPU clusters, ML platforms, or distributed serving infrastructure.

  • Built automation that reduced toil, improved recovery time, or prevented repeat incidents.

  • Led incident response for severe outages with clear communication across engineering and leadership.

  • Created practical SLOs, dashboards, alerts, runbooks, or release gates that improved production reliability.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: We offers generous health, dental, and vision benefits as well as 401(k) company match.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
619,872 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$60k – $109k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$44k – $101k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$51k – $117k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$35k – $79k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$23k – $54k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Frontend
Lighthouse
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 3 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 3 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$115k – $180k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • Seattle • Raleigh • New York
Apply
$97k – $124k per year • In office • Full-Time • 1+ year exp • San Francisco
Design
Canva
Management
Slack
Google Workspace
Apply
$200k – $240k per year • Remote/Hybrid • San Francisco
AI/ML
AI Agents
LLM
RAG
Apply
$152k – $301k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin • San Francisco • Fort Worth • Los Angeles
Marketing
Salesforce
Apply
$285k – $335k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
DevOps
VMWare
containerd
Kubernetes
KVM
QEMU
Hyper-V
Apply
See all jobs
This is one of many
619,872 more open roles from verified company boards, updated every day.