368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$200k – $420k per year
Location
In office (Palo Alto)
Seniority
Staff · 5+ years exp
Overview
Company
Impact
Profile match
River AI is an artificial intelligence company headquartered in Palo Alto, California, and founded in 2026 by xAI co-founder Igor Babuschkin. The company is building what it calls an open AI stack, spanning personal AI models, tooling for developers to train and serve their own models, and a custom system on chip with an onboard machine learning accelerator. It raised 1.1 billion dollars led by General Catalyst with strategic investment from NVIDIA and AMD Ventures.

At River, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, custom training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional performance and kernel generation engineers to build the foundational compute engine for our high-performance custom silicon. In this role, you will design and implement robust kernel generators that programmatically emit optimized low-level assembly code for our greenfield hardware architecture.

You will bridge the gap between high-level compilation and raw hardware capability, pushing our custom architecture to its absolute theoretical limits for critical deep learning operations (including GEMMs, FlashAttention, and custom activations). You will collaborate closely up and down the stack with compiler engineers, silicon architects, and deep learning researchers to unlock maximum compute efficiency.

What You’ll Do

  • Kernel Generator Development: Design and build C++ code-generation frameworks and meta-programming toolchains that automatically emit optimized custom ISA assembly code.
  • Low-Level Compute Optimization: Author and optimize core deep learning primitives (GEMM/MatMul, Attention mechanisms, Convolutions, and element-wise layers) directly targeted at our custom hardware.
  • Microarchitectural Tuning: Hand-craft and automate instruction scheduling, register allocation, and software pipelining to maximize ALU utilization and hide execution latency on our silicon.
  • Memory Hierarchy Management: Design sophisticated tiling, double-buffering, and data-movement strategies to optimize on-chip SRAM utilization and minimize memory bandwidth bottlenecks.
  • HW/SW Co-Design: Partner with the RTL and architecture teams to evaluate hardware simulations, provide feedback on the ISA, and influence the design of future compute units based on kernel execution profiles.
  • Performance Profiling & Validation: Benchmark generated assembly against hardware simulators and silicon, utilizing hardware performance counters to eliminate performance gaps and ensure mathematical correctness.

Minimum Qualifications:

  • Bachelor’s degree in Computer Engineering, Computer Science, Electrical Engineering, or a related field, and 5+ years of practical industry experience in low-level performance programming.
  • Deep understanding of hardware programming models (e.g., CUDA, Triton, CUTLASS, or custom accelerator assembly) and a proven track record of shipping highly optimized kernels.
  • Advanced knowledge of Computer Architecture, including vector units, execution pipelines, register files, and complex memory hierarchies (caches, SRAM, HBM/DRAM).
  • Proficiency in modern C++ for building robust, scalable meta-programming and code-generation frameworks.
  • Strong mathematical foundation in linear algebra operations and deep learning primitives.
  • A highly collaborative mindset to push boundaries and co-design effectively with hardware and compiler teams.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Deep familiarity with implementing microarchitectural optimizations for Tensor Cores, matrix multiply-accumulate units, or custom vector extensions.
  • Experience utilizing advanced C++ template metaprogramming or code-generation techniques to automate the creation of heavily parameterized kernel variants.
  • Advanced experience with low-level hardware profiling tools, execution tracing, and utilizing performance counters to identify cache misses, pipeline stalls, and ALU bubbles.

Logistics

  • Location: This role is based in Austin, Texas or Palo Alto, California.
  • Compensation: Depending on background, skills, experience, and location, the expected annual salary range for this position is $200,000 - $420,000 USD.
  • Visa Sponsorship: We sponsor visas. We can't guarantee success for every candidate or role, but if you're the right fit, we're committed to working through the visa process.
  • Benefits: River AI offers generous health, dental, and vision benefits, unlimited PTO, and relocation support as needed.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$118k – $240k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Tysons
C#
Go
JavaScript
Python
SQL
TypeScript
C#
.NET
AI/ML
Embeddings
Human-in-the-Loop
LLM Guardrails
RAG
Semantic Search
Semantic Search
DevOps
AWS
CI/CD
Vector
Cybersecurity
Least Privilege
Apply
$113k – $188k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Tysons
Python
SQL
Python
pySpark
Databases
Amazon Redshift
Apache Iceberg
Databricks
Delta Lake
AI/ML
Airflow
Anomaly Detection
ChatGPT
Copilot
Cursor
GraphRAG
Knowledge Graph
RAG
Spark
DevOps
Amazon Kinesis
Amazon S3
AWS
AWS CDK
AWS Lambda
AWS Step Functions
CI/CD
CloudFormation
Docker
Git
GitHub
GitHub Actions
IAM
Jenkins
Terraform
Vector
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Power BI
Tableau
Apply
$71k – $112k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austria
C#
C++
Python
Visual Basic
Apply
$150k per year • In office • Full-Time • New York
C++
Python
Apply
In office • Full-Time • Bachelor's Degree • Gurgaon
C++
COBOL
Java
SQL
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
Python
Rust
TypeScript
JavaScript
AI/ML
LLM
Pre-training
Structured Outputs
Function Calling
Frontend
React.js
Mobile
React Native
DevOps
CI/CD
Kubernetes
Terraform
Apply
$200k – $420k per year • In office • 5+ years exp • Bachelor's Degree • Palo Alto
C++
C++
LLVM
PyTorch C++
AI/ML
CUDA Toolkit
PyTorch
Pre-training
Apply
$200k – $420k per year • In office • 5+ years exp • Bachelor's Degree • Palo Alto
C++
SystemC
C
C++
LLVM
C
GCC
AI/ML
CUDA Toolkit
Pre-training
DevOps
QEMU
Apply
$131k – $333k per year (Estimated) • In office • Palo Alto
AI/ML
Pre-training
Apply
$200k – $420k per year • In office • Bachelor's Degree • Palo Alto
AI/ML
Diffusion Models
Fine-tuning
JAX
Multimodal AI
PEFT
PyTorch
Reinforcement Learning
RLHF
Transformers
Edge AI
Pre-training
Apply
$110k – $240k per year (Estimated) • In office • Bachelor's Degree • Palo Alto
Java
Python
Scala
AI/ML
AI Agents
Fine-tuning
LLM Guardrails
DevOps
AWS
Azure
GCP
Git
GitHub
Marketing
Salesforce
Apply
Chief of Staff 9 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
$139k – $294k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
$137k – $292k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.