412,983open jobs
14,312companies
71,936added this week
Browse all
Salary
$186k – $383k per year (Estimated)
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Index : kitd this commitmain process supervisor for OpenBSDabout summary refs log tree commit diff log msgauthorcommitterrange KITD(8) System Manager's Manual KITD(8) NAME kitd — process supervisor SYNOPSIS kitd [-d] [-c cooloff] [-m maximum] [-n name] [-t restart] command ...

Our mission is general causal intelligence; AI that is capable of (1) predicting the future and (2) identifying the actions to alter it.

To achieve this breakthrough, we are building a Large Physics foundation Model (LPM) because physical systems, unlike text or images, are governed by verifiable cause and effect. We believe that scaling on physics will enable an understanding of causality required to predict and control physical systems, starting with weather.

Our founding team has built and deployed AI against the physical world in robotics, drug discovery, and particle physics at institutions like DeepMind, Waymo, Cruise, Insitro, Nabla Bio, and CERN.

We look for infrastructure engineers who are excited to tackle unsolved problems. Everything we do - training, evaluation, serving - runs on our GPU fleet. Your mission is to design, build, and operate the supercomputing environment underneath it all, delivering performant, reliable, and cost-efficient compute to ensure research is able to iterate rapidly at scale.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning

  • Extend scheduling and orchestration systems (e.g. Kubernetes, Slurm) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads

  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers

  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage

  • Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do

  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs

What we're looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.

  • Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)

  • Strong systems background: Linux, networking, storage, infrastructure-as-code

  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings

  • Understanding of monitoring, logging, observability, and version control best practices for ML systems

  • Familiarity with CUDA/NCCL and performance profiling for distributed workloads

  • Owns deliverables end-to-end, from requirements through autonomous execution

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
412,983 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$22k – $52k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Malaysia
DevOps
GCP
Azure
AWS
Apply
Senior Java Engineer 2 hours ago
$66k – $142k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Madrid
Python
Java
PowerShell
Java
Spring Boot
DevOps
Terraform
Ansible
GCP
OpenShift
Helm
Azure
CI/CD
AWS
Kubernetes
Configuration Management
Apply
$20k – $50k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Java
TypeScript
Node JS
COBOL
Python
Poetry
Java
Maven
Spring Boot
Spring MVC
Gradle
COBOL
IBM MQ
Databases
PostgreSQL
RabbitMQ
ActiveMQ
Apache Kafka
Kafka
Frontend
GraphQL
Angular
React.js
Mobile
JUnit
DevOps
gRPC
GCP
OpenShift
WebSockets
Azure
CI/CD
AWS
Kubernetes
QA
Pytest
Apply
$32k – $77k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Buenos Aires
SQL
AI/ML
Edge AI
DevOps
GCP
Analytics
Power BI
Apply
$54k – $124k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Quebec
DevOps
GCP
Apply
$182k – $374k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
DevOps
Terraform
GCP
Pulumi
SLURM
Azure
AWS
Docker
Kubernetes
IAM
Cybersecurity
Okta
SOC 2
FedRAMP
Zero Trust
Threat Modeling
Microsoft Entra ID
Apply
$209k – $429k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
Reinforcement Learning
Post-training
Robotics
Reinforcement Learning
Apply
$197k – $404k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
Multimodal AI
Computer Vision
World Models
Robotics
Sensor Fusion
Apply
$202k – $415k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
Interpretability
Apply
$184k – $377k per year (Estimated) • In office • Full-Time • San Francisco
DevOps
Terraform
GCP
Azure
AWS
Docker
Kubernetes
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
Fullstack Engineer 4 hours ago
$100k – $150k per year • Equity 0.2–2% • Remote • Full-Time • 1+ year exp • San Francisco
Python
JavaScript
C++
DevOps
AWS
Apply
See all jobs
This is one of many
412,983 more open roles from verified company boards, updated every day.