435,295open jobs
15,159companies
67,452added this week
Browse all
Salary
$215k – $325k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thunder Compute provides virtualised graphics processing units for machine learning developers. Its technology shares accelerators across workloads to cut cost. The company targets small teams and researchers.

Company

The world is building massive amounts of GPU capacity. Meanwhile, deployed GPUs are only 20% utilized.

This is because GPUs are not virtualized, while every other type of hardware is. For example CPUs and storage are allocated through virtual abstractions which efficiently manage the physical hardware, while GPUs are statically allocated on a one-to-one basis.

Thunder Compute is building this virtualization layer for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic.

Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer.

We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.

Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center.

This enables something like “Ceph for GPUs”: GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.

Role

Your work will focus on building the core C++ systems behind our virtualization layer. This includes low-latency performance optimization, distributed systems debugging, production reliability, and research into new techniques for improving GPU utilization.

You will take ownership of complex systems from early experimentation through production deployment. Example projects may include:

  • Profiling and reducing latency across remote CUDA operations

  • Building high-performance networking and data-transfer paths

  • Debugging failures across customer processes, our userspace runtime, the network, and remote GPU servers

  • Improving support for process forking, signals, multithreading, dynamic linking, and unusual application behavior

  • Designing systems for GPU allocation, scheduling, failure recovery, and observability

  • Researching and productionizing new GPU virtualization and oversubscription techniques

  • Expanding compatibility across CUDA applications, frameworks, and GPU architectures

You will spend your days bouncing between the weeds of complex, performance-critical systems that are live in production. One week, you may be tracing a synchronization bug across a distributed CUDA workload; the next, you may be redesigning a hot data path to remove microseconds of overhead.

This work is not easy. It blends the hardest parts of systems research and production engineering.

We look for exceptional low-level engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable systems code.

Core Technical Skills

  • Exceptional modern C++ ability, including memory management, concurrency, performance optimization, and systems-level abstraction design

  • Deep understanding of operating systems, low-level networking, compilers, distributed systems, or computer architecture

  • Experience building and operating performance-critical C++ systems in production

  • Strong Linux systems programming and debugging ability

  • Ability to reason through unfamiliar systems across multiple layers of the stack

Must Haves

  • Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines

  • Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight

  • Strong ownership over correctness, reliability, performance, and operational outcomes

  • Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner

  • Willingness to work directly with customers and investigate difficult production failures

Preferred

  • Experience with CUDA, GPU systems, compilers, runtime interception, dynamic linking, high-performance networking, or distributed computing

  • Experience at a trading firm such as Citadel Securities or Jane Street; a hardware or AI infrastructure company such as NVIDIA or SambaNova; a systems research group; or a similarly demanding engineering environment

  • Strong computer science fundamentals demonstrated through academic work, systems research, competitive programming, open-source contributions, or exceptional professional experience

  • Experience taking new systems research from a paper or prototype into a reliable production system

Why Join

You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company.

You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The systems you build will form the foundation of a new infrastructure layer for GPU computing.

Logistics

  • You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities

  • This role is full-time and in person, five days per week, at our office in downtown San Francisco

  • Relocation support and visa sponsorship are available

Benefits

  • Competitive salary and meaningful equity

  • Daily lunch, snacks, and coffee

  • Team dinners and events

  • 401(k)

  • Health, dental, and vision insurance

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,295 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$93k – $233k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Triton Inference Server
Embeddings
Quantization
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
RAG
Reranking
NVIDIA NIM
CUDA
Triton
Hugging Face
NCCL
NVLink
cuDNN
Speculative Decoding
KV Cache
Agentic Workflows
Tool Use
DevOps
CI/CD
GitOps
Kubernetes
Platform Engineering
Apply
$33k – $87k per year (Estimated) • Remote • Full-Time
JavaScript
TypeScript
Node JS
Node JS
BullMQ
Databases
PostgreSQL
AI/ML
Cursor
Claude
Claude Code
Prompt Engineering
AI Agents
AWS Bedrock
LLM
OpenAI
Anthropic
Tool Use
Frontend
Next.js
React.js
React Query
DevOps
Terraform
AWS
Amazon ECS
Amazon CloudWatch
QA
Sentry
Apply
Remote/Hybrid • 5+ years exp
Python
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
pgvector
AI/ML
Claude
Claude Code
Embeddings
Prompt Engineering
Function Calling
LLM
RAG
Semantic Search
Anthropic
Structured Outputs
Semantic Search
Tool Use
Frontend
React.js
DevOps
Azure
Apply
$52k – $129k per year (Estimated) • Remote • Full-Time • 5+ years exp
Python
JavaScript
TypeScript
SQL
Python
FastAPI
Databases
PostgreSQL
Neo4j
Apache Iceberg
Delta Lake
TimescaleDB
MinIO
InfluxDB
Apache Kafka
Trino
AI/ML
LangGraph
Polars
LangChain
Spark
Airflow
Model Context Protocol
AI Agents
LangSmith
LiteLLM
Ollama
Ragas
Flink
LLM
RAG
Semantic Search
Hybrid Search
Time Series Forecasting
OpenAI
Anthropic
Semantic Search
Knowledge Graph
Frontend
Angular
React.js
DevOps
Rest API
Terraform
Azure
Kubernetes
Vector
IoT
MQTT
OPC UA
Apply
In office • 10+ years exp • Bachelor's Degree
Python
Go
JavaScript
TypeScript
C++
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Embeddings
Prompt Engineering
Function Calling
AI Agents
Semantic Kernel
CrewAI
LLM
RAG
Google ADK
Hallucination
OpenAI
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Vector
Apply
Chief of Staff 2 months ago
$150k – $225k per year • In office • Full-Time • San Francisco
AI/ML
Anthropic
CoreWeave
DevOps
VMWare
Apply
$150k – $250k per year • In office • Full-Time • San Francisco
Python
Go
JavaScript
TypeScript
AI/ML
CUDA Toolkit
CUDA
Anthropic
CoreWeave
Frontend
Next.js
React.js
DevOps
gRPC
AWS
Kubernetes
AWS Lambda
Apply
$76k – $92k per year • Remote • Internship • Bachelor's Degree • San Francisco
SQL
Analytics
SSIS
SSAS
Management
Microsoft Project
Apply
$70k – $82k per year • Remote • Internship • San Francisco
Analytics
Microsoft Excel
Apply
$70k – $82k per year • In office • Internship • San Francisco
Apply
$79k – $95k per year • Remote • Full-Time • San Francisco
Apply
$213k – $374k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Chicago • New York • Atlanta • San Francisco
AI/ML
AI Agents
Agentforce
Agentic Workflows
Marketing
Salesforce
Apply
See all jobs
This is one of many
435,295 more open roles from verified company boards, updated every day.