569,038open jobs
23,263companies
77,895added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Founded by the creators and core maintainers of vLLM, we sit at the intersection of models and hardware-a position that took years to build.

About the Role

We're looking for a hands-on cluster administration engineer to own and operate the high-performance GPU compute infrastructure that keeps Inferact engineering productive. Inferact runs on expensive, high-performance GPU and HPC clusters across neo-cloud and dedicated compute providers. Your job is to make sure that infrastructure is healthy, available, observable, and usable around the clock.

You'll take ownership of cluster health, GPU availability, monitoring, alerting, scheduling, access, diagnostics, and incident response across the systems our engineers rely on every day. You'll work closely with engineering leadership and infrastructure owners to standardize how we provision, operate, debug, and scale compute across providers. Your work will directly impact how fast Inferact can build, test, and improve the systems powering vLLM.

Skills and Qualifications

Minimum qualifications:

  • Bachelor's degree or equivalent experience in computer science, engineering, systems administration, or similar.

  • Hands-on experience administering large compute clusters, HPC environments, university or research clusters, supercomputing systems, or production GPU clusters.

  • Strong Linux systems administration fundamentals across networking, processes, storage, package management, shell scripting, logs, access control, and system debugging.

  • Experience operating GPU servers, including driver management, GPU health monitoring, node failures, memory errors, scheduler issues, and hardware diagnostics.

  • Experience with cluster scheduling and resource allocation using SLURM, Kubernetes, or equivalent tooling.

  • Ability to own urgent infrastructure incidents end-to-end when compute issues are blocking engineering teams.

  • Ability to automate operational workflows using Bash, Python, Ansible, Terraform, Helm, or similar tooling.

Preferred qualifications:

  • Experience operating GPU compute across providers such as Lambda, CoreWeave, Crusoe, Nebius, Together, Fireworks, RunPod, or similar environments.

  • Experience improving cluster utilization, reducing idle or unavailable GPU capacity, and debugging scheduling or resource contention issues.

  • Familiarity with high-performance GPU networking such as InfiniBand, RoCE, NVLink / NVSwitch, RDMA, NCCL, or equivalent systems.

  • Experience with storage for HPC or ML workloads, including NFS, Lustre, Ceph, distributed filesystems, or other high-throughput storage systems.

  • Experience managing secure access, identity, permissions, SSH, VPNs, bastion hosts, secrets, and basic infrastructure security hygiene.

  • Background in research computing, scientific computing, ML infrastructure, SRE, platform engineering, or infrastructure operations for engineering-heavy teams.

Bonus points if you have:

  • Managed GPU or HPC infrastructure in a university lab, national lab, research institution, AI infrastructure company, hedge fund, HFT firm, or large-scale ML platform team.

  • Built monitoring, alerting, runbooks, health checks, or remediation workflows that materially reduced operational toil or incident resolution time.

  • Operated Kubernetes clusters for ML or GPU workloads at meaningful scale.

  • Standardized provisioning, diagnostics, monitoring, and operating patterns across multiple compute providers.

  • Carried real operational responsibility for infrastructure used by many engineers or researchers.

Logistics

  • Location: This role is based in San Francisco, California. Will consider remote in the US for exceptional candidates.

  • Compensation: Depending on background, skills, and experience, the expected annual salary range for this position is $200,000 - $400,000 USD + equity.

  • Visa sponsorship: We sponsor visas on a case-by-case basis.

  • Benefits: Inferact offers generous health, dental, and vision benefits as well as 401(k) company match.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
569,038 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$13k – $28k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Go
Databases
InfluxDB
DevOps
gRPC
Ansible
OpenTelemetry
containerd
Prometheus
CI/CD
Git
Docker
Kubernetes
Ubuntu
Grafana
Platform Engineering
KVM
Telegraf
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Ottawa
Python
JavaScript
C++
Management
Agile
Apply
UX Researcher 4 hours ago
$64k – $147k per year (Estimated) • In office • Full-Time • PhD • United Kingdom
Python
SQL
SPSS
Apply
$14k – $36k per year (Estimated) • In office • Tashkent
Python
DevOps
Zabbix
Grafana
Nagios
Apply
$33k – $58k per year (Estimated) • Remote • Moscow
Python
Go
JavaScript
TypeScript
C++
C++
STL
DevOps
Git
Apply
$125k – $170k per year • In office • Full-Time • San Francisco
AI/ML
vLLM
Apply
$165k – $355k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Python
Go
Rust
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
JAX
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
PyTorch
Ray
Mixture of Experts
CUDA
Triton
TPU
NCCL
InfiniBand
ROCm
MLIR
XLA
CUTLASS
KV Cache
DevOps
Terraform
Helm
SLURM
Kubernetes
Apply
Head of Legal 2 days ago
$194k – $373k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
AI/ML
vLLM
LLM
Apply
HR / People Lead 2 days ago
$180k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
vLLM
Apply
$163k – $336k per year (Estimated) • Remote/Hybrid • Full-Time
Python
AI/ML
vLLM
Multimodal AI
Diffusion Models
AI Agents
SGLang
TensorRT
LLaMA-Factory
TensorRT-LLM
TGI
Unsloth
PyTorch
LLM
Mixture of Experts
KV Cache
Apply
$106k – $130k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Chicago • Scottsdale
DevOps
GCP
Azure
CI/CD
AWS
Harbor
Platform Engineering
Incident Management
Apply
$186k – $345k per year (Estimated) • In office • Full-Time • 12+ years exp • Master's Degree • San Francisco
DevOps
Incident Management
Apply
$160k – $236k per year • Equity • Remote/Hybrid • Full-Time • 2+ years exp • Chicago • Austin • San Francisco
Apply
$63k – $122k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Chicago • San Francisco • Dallas • Boston • Philadelphia
Python
JavaScript
Java
TypeScript
SQL
C++
AI/ML
Prompt Engineering
Function Calling
AI Agents
LLM
Anthropic
Agentic Workflows
Tool Use
Frontend
Angular
React.js
DevOps
GCP
Azure
AWS
GitHub
Cybersecurity
Threat Modeling
Apply
$40k – $100k per year • In office • Full-Time • 1+ year exp • San Francisco
AI/ML
Claude
AI Agents
DevOps
Azure
Analytics
A/B Testing
Design
Adobe Photoshop
Figma
Marketing
Google Ads
Apply
See all jobs
This is one of many
569,038 more open roles from verified company boards, updated every day.