570,882open jobs
23,625companies
78,041added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

About the Role

We're looking for an cloud orchestration engineer to build the operational backbone that keeps vLLM running reliably at massive scale. You'll design the systems for cluster management, deployment automation, and production monitoring that enable teams worldwide to serve AI models without friction. You'll ensure that vLLM deployments are observable, debuggable, and recoverable, turning operational complexity into infrastructure that just works.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, or similar.

  • Strong experience with Kubernetes and container orchestration at scale.

  • Experience designing and implementing custom Kubernetes operators.

  • Proficiency in Python/Rust/Go and infrastructure-as-code tools (Terraform, Helm, etc).

  • Experience managing GPU clusters and debugging hardware issues.

  • Ability to work across cloud platforms (AWS, GCP, Azure) and on-premise infrastructure.

Preferred qualifications:

  • Experience with ML-specific orchestration tools (Ray, Slurm).

  • Knowledge of GPU scheduling, multi-tenancy, and resource optimization.

  • Familiarity with vLLM deployment patterns and configuration.

  • Track record of improving operational reliability for ML systems.

Bonus points if you have:

  • Experience deploying inference systems on large-scale GPU (1,000+) clusters.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
570,882 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$11k – $27k per year (Estimated) • In office • 2+ years exp • Bengaluru
Python
SQL
AI/ML
Prompt Engineering
RAG
Knowledge Graph
DevOps
GCP
OpenShift
New Relic
OpenTelemetry
Dynatrace
Azure
AWS
Kubernetes
AIOps
Incident Management
SLI/SLO/SLA
Amazon CloudWatch
Management
Jira
ServiceNow
QA
Postman
Apply
In office • Full-Time • 3+ years exp • Chennai
Python
Bash
Databases
Apache Kafka
DevOps
Rest API
CI/CD
AWS
Analytics
Apache NiFi
Apply
DevOps Engineer 6 min ago
In office • Full-Time • 5+ years exp • Chennai
Python
Bash
Databases
Apache Kafka
DevOps
Rest API
CI/CD
AWS
Analytics
Apache NiFi
Apply
In office • Full-Time • London
DevOps
GCP
Azure
AWS
Apply
$82k – $137k per year • In office • Full-Time • Master's Degree • United States
Python
MATLAB
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 2 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$106k – $130k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Chicago • Scottsdale
DevOps
GCP
Azure
CI/CD
AWS
Harbor
Platform Engineering
Incident Management
Apply
$186k – $345k per year (Estimated) • In office • Full-Time • 12+ years exp • Master's Degree • San Francisco
DevOps
Incident Management
Apply
$160k – $236k per year • Equity • Remote/Hybrid • Full-Time • 2+ years exp • Chicago • Austin • San Francisco
Apply
$40k – $100k per year • In office • Full-Time • 1+ year exp • San Francisco
AI/ML
Claude
AI Agents
DevOps
Azure
Analytics
A/B Testing
Design
Adobe Photoshop
Figma
Marketing
Google Ads
Apply
$74k – $219k per year • Remote/Hybrid • Full-Time • 3+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Management
Agile
Apply
See all jobs
This is one of many
570,882 more open roles from verified company boards, updated every day.