435,297open jobs
15,143companies
67,158added this week
Browse all
Salary
$150k – $250k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thunder Compute provides virtualised graphics processing units for machine learning developers. Its technology shares accelerators across workloads to cut cost. The company targets small teams and researchers.

Company

The world is building massive amounts of GPU capacity. Meanwhile, deployed GPUs are only 20% utilized.

This is because GPUs are not virtualized, while every other type of hardware is. For example CPUs and storage are allocated through virtual abstractions which efficiently manage the physical hardware, while GPUs are statically allocated on a one-to-one basis.

Thunder Compute is building this virtualization layer for GPUs. We have raised over $17.5M from Matrix Partners, Y Combinator, and leading angels from Coreweave, Microsoft, Cognition, and Anthropic.

Leading solutions for underutilization sit at the workload layer and are therefore only able to optimize specific use cases. We believe the ideal cluster optimization solution must be invisible to developers and compatible with all workloads; hence, it must sit at the systems layer.

We are a team of systems researchers productionizing cutting-edge GPU virtualization research to build this general-purpose optimization layer.

Concretely, our virtualization library abstracts GPUs across TCP networking. We use a userspace shim library, loaded through LD_PRELOAD, to intercept CUDA calls and send them over gRPC to a host server connected to a physical GPU elsewhere in the data center.

This enables something like “Ceph for GPUs”: GPUs become network resources that can be abstracted, pooled, and dynamically allocated across a cluster to improve utilization without requiring developers to modify their workloads.

Role

Your work will focus on building the cloud infrastructure surrounding our GPU virtualization layer. This includes the Go backbone of our cloud platform, Kubernetes-based orchestration, production reliability, networking, storage, billing infrastructure, and the systems used to deploy and operate GPU capacity at scale.

You will take ownership of complex infrastructure from early design through production deployment. Example projects may include:

  • Building control-plane services for provisioning and managing virtual GPU instances

  • Designing reliable systems for GPU allocation, scheduling, and lifecycle management

  • Improving our unconventional Kubernetes deployment, which acts as a form of hypervisor for customer workloads

  • Building infrastructure for networking, storage, authentication, billing, and usage metering

  • Automating the deployment and operation of GPU hosts across cloud providers and customer data centers

  • Debugging failures across customer workloads, Kubernetes, our control plane, the network, and physical GPU infrastructure

  • Designing systems for failure recovery, capacity management, observability, and incident response

  • Improving the security, reliability, and operational simplicity of the platform as it scales

  • Working directly with customers to diagnose problems and deploy Thunder Compute in new environments

You will spend your days bouncing between the weeds of complicated production infrastructure that is live and used by customers. One week, you may be debugging a networking failure across a Kubernetes cluster; the next, you may be redesigning the provisioning system to make deployments faster and more reliable.

This work is not easy. It blends the hardest parts of cloud infrastructure, distributed systems, and production engineering.

We look for exceptional engineering talent, strong work ethic, and extreme attention to detail. We must move quickly while shipping high-quality, reliable infrastructure.

Core Technical Skills

  • Exceptional Go ability, including concurrency, distributed systems design, API design, and production service development

  • Deep understanding of Kubernetes, containers, Linux, networking, storage, or cloud infrastructure

  • Experience building and operating critical production systems

  • Strong systems debugging and operational ability

  • Ability to reason through unfamiliar systems across multiple layers of the stack

  • Working knowledge of Python; familiarity with TypeScript or Next.js is helpful

Must Haves

  • Strong work ethic and the ability to independently push a project from an experimental prototype through 100% completion under tight deadlines

  • Attention to detail and the ability to deliver production-ready, thoroughly tested code without significant oversight

  • Strong ownership over correctness, reliability, performance, and operational outcomes

  • Ability to debug ambiguous problems without a clear reproduction, existing playbook, or obvious owner

  • Willingness to work directly with customers and investigate difficult production failures

  • Strong communication skills and the ability to coordinate across engineering, customers, and external infrastructure providers

Preferred

  • Experience with Kubernetes internals, container runtimes, cloud networking, distributed storage, infrastructure security, or large-scale control planes

  • Experience building high-stakes production infrastructure at a trading firm such as Citadel Securities or Jane Street; a cloud provider such as AWS, CoreWeave, or Lambda; an AI infrastructure company; or a similarly demanding engineering environment

  • Strong computer science fundamentals demonstrated through academic work, distributed systems research, open-source contributions, or exceptional professional experience

  • Experience designing and operating infrastructure across multiple cloud providers or on-premise environments

  • Experience taking a new infrastructure system from an early prototype into a reliable production platform

Why Join

You will join early enough to meaningfully shape the architecture, engineering standards, and technical direction of the company.

You will work directly with the founders on a category-defining systems problem, with a short path between writing code and seeing it run in production. The infrastructure you build will operate a new foundational layer for GPU computing.

Unlike at a large company, you will not be restricted to one small component of a much larger system. You will own broad, technically difficult areas of the platform and have the opportunity to grow into senior technical and engineering leadership as the company scales.

Logistics

  • You will report to co-founder and CTO Brian Model, formerly a Quantitative Developer at Citadel Securities

  • This role is full-time and in person, five days per week, at our office in downtown San Francisco

  • Relocation support and visa sponsorship are available

Benefits

  • Competitive salary and meaningful equity

  • Daily lunch, snacks, and coffee

  • Team dinners and events

  • 401(k)

  • Health, dental, and vision insurance

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,297 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$141k – $212k per year • Remote/Hybrid • Full-Time • Tampa • Irving
Python
Go
AI/ML
LangChain
Claude
LlamaIndex
MLFlow
Vertex AI
Prompt Engineering
AI Agents
Kubeflow
Gemini
LLM
RAG
Anomaly Detection
Amazon SageMaker
GPT-4
DevOps
OpenShift
GitHub Actions
Loki
OpenTelemetry
Prometheus
CI/CD
ArgoCD
Kubernetes
Grafana
Platform Engineering
Tekton
Progressive Delivery
GitHub
Apply
$87k – $131k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Jacksonville • Irving • San Antonio • Tampa
Python
SQL
SAS
Analytics
Tableau
Apply
$71k – $182k per year (Estimated) • In office • Internship • Bachelor's Degree • Indianapolis
Python
Apply
$50k – $52k per year • In office • Internship • Bachelor's Degree • Carmel
Python
PowerShell
DevOps
Terraform
Azure DevOps
Azure
CI/CD
Git
Apply
In office • Full-Time • PhD • Birmingham
JavaScript
Frontend
Parcel
Analytics
Microsoft Excel
Apply
Chief of Staff 2 months ago
$150k – $225k per year • In office • Full-Time • San Francisco
AI/ML
Anthropic
CoreWeave
DevOps
VMWare
Apply
$215k – $325k per year • In office • Full-Time • San Francisco
C++
AI/ML
CUDA Toolkit
CUDA
Anthropic
CoreWeave
DevOps
gRPC
Apply
$76k – $92k per year • Remote • Internship • Bachelor's Degree • San Francisco
SQL
Analytics
SSIS
SSAS
Management
Microsoft Project
Apply
$70k – $82k per year • Remote • Internship • San Francisco
Analytics
Microsoft Excel
Apply
$70k – $82k per year • In office • Internship • San Francisco
Apply
$79k – $95k per year • Remote • Full-Time • San Francisco
Apply
$213k – $374k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Chicago • New York • Atlanta • San Francisco
AI/ML
AI Agents
Agentforce
Agentic Workflows
Marketing
Salesforce
Apply
See all jobs
This is one of many
435,297 more open roles from verified company boards, updated every day.