597,814open jobs
30,191companies
86,075added this week
Browse all
Salary
$200k – $420k per year
Location
In office (Palo Alto)
Overview
Company
Impact
Profile match
River AI is an artificial intelligence company headquartered in Palo Alto, California, and founded in 2026 by xAI co-founder Igor Babuschkin. The company is building what it calls an open AI stack, spanning personal AI models, tooling for developers to train and serve their own models, and a custom system on chip with an onboard machine learning accelerator. It raised 1.1 billion dollars led by General Catalyst with strategic investment from NVIDIA and AMD Ventures.

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional inference systems engineers to build the engines that serve large models through the River API. Your goal is to deliver fast, reliable inference while making efficient use of GPU compute and memory.

You will take ownership of the serving runtime, from request scheduling and continuous batching to KV-cache management, distributed model execution, and checkpoint loading. Your work will support both customer-facing inference and the sampling workloads that power reinforcement learning.

Working closely with GPU kernel engineers, researchers, and infrastructure engineers, you will bring new models into production and improve their performance across realistic workloads. You will measure success through latency, throughput, reliability, and cost, with careful attention to numerical correctness and model behavior.

What You’ll Do

  • Optimize inference for dense and mixture-of-experts models, including fine-tuned models and adapters.
  • Improve batching, caching, and admission control to balance throughput, latency, memory use, and fairness.
  • Accelerate multi-GPU execution, communication, and model loading while preserving model-version consistency.
  • Improve RL sampling throughput while keeping samples and log probabilities tied to the correct model version.
  • Build reliable streaming, cancellation, and recovery under failures and overload.
  • Profile bottlenecks and validate improvements through reproducible performance and correctness tests.

Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience building inference engines or performance-sensitive distributed services.
  • Strong understanding of transformer inference, GPU memory, concurrency, and networking.
  • Proficiency in Python and C++ or Rust.
  • Strong debugging and profiling skills across models, runtimes, and services.
  • A collaborative mindset and strong ownership of engineering outcomes.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Experience extending SGLang, vLLM, TensorRT-LLM, or similar frameworks.
  • Work on advanced serving techniques, such as speculative decoding or disaggregated prefill and decode.
  • Familiarity with expert parallelism, GPU collectives, and quantized inference.
  • Experience with multi-adapter serving, dynamic checkpoint loading, or RL sampling.
  • Familiarity with CUDA graphs, custom kernels, and NVIDIA profiling tools.
  • Experience operating model-serving systems under production traffic.

Logistics & Benefits

  • Location: Palo Alto, California.
  • Compensation: $200,000-$420,000 USD annual base pay, depending on experience and skills.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and support the process for the right candidate.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
597,814 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$165k – $199k per year • In office • Full-Time • 8+ years exp • El Segundo
Python
C++
Apply
$158k – $189k per year • In office • Full-Time • 8+ years exp • Austin
Python
C++
Apply
$158k – $189k per year • In office • Full-Time • 8+ years exp • Westminster
Python
C++
Apply
$20k – $57k per year (Estimated) • In office • Internship • Chennai
JavaScript
Java
SQL
C++
Java
Apache Tomcat
Databases
MySQL
Oracle
DevOps
Rest API
Management
Agile
Apply
$19k – $53k per year (Estimated) • In office • Internship • 3+ years exp • Chennai
Python
SQL
PowerShell
Databases
Redis
Azure SQL Database
DevOps
Splunk
Terraform
Ansible
Helm
New Relic
Rancher
Azure
Docker
Kubernetes
Incident Management
SLI/SLO/SLA
Cybersecurity
GDPR
Management
Agile
Scrum
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Fine-tuning
SGLang
Accelerate
PyTorch
Mixture of Experts
CUDA
Triton
Pre-training
CUTLASS
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
LoRA
CUDA Toolkit
Fine-tuning
Reinforcement Learning
JAX
PEFT
PyTorch
Mixture of Experts
CUDA
Pre-training
NCCL
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
JavaScript
Rust
TypeScript
AI/ML
Function Calling
LLM
Pre-training
Structured Outputs
Tool Use
Frontend
React.js
Mobile
React Native
DevOps
Terraform
CI/CD
Kubernetes
Apply
Office Manager 1 day ago
$60k – $120k per year • In office • San Francisco
AI/ML
Post-training
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
AI/ML
Fine-tuning
RLHF
Reinforcement Learning
JAX
Multimodal AI
Diffusion Models
PEFT
Transformers
PyTorch
Pre-training
Edge AI
Apply
$98k – $209k per year (Estimated) • In office • Full-Time • 3+ years exp • Palo Alto
Python
TypeScript
Bash
DevOps
Terraform
Ansible
GCP
GitHub Actions
CloudFormation
Pulumi
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Fine-tuning
SGLang
Accelerate
PyTorch
Mixture of Experts
CUDA
Triton
Pre-training
CUTLASS
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
Rust
C++
C++
PyTorch C++
AI/ML
LoRA
CUDA Toolkit
Fine-tuning
Reinforcement Learning
JAX
PEFT
PyTorch
Mixture of Experts
CUDA
Pre-training
NCCL
Apply
$60k – $140k per year • Equity • In office • Full-Time • 8+ years exp • PhD • Palo Alto
Apply
$60k – $140k per year • Equity • In office • Full-Time • PhD • Palo Alto
Python
C
C++
Fortran
C
MPI
AI/ML
CUDA Toolkit
CUDA
ROCm
DevOps
HPC
Apply
See all jobs
This is one of many
597,814 more open roles from verified company boards, updated every day.