721,262open jobs
43,022companies
101,631added this week
Browse all
Salary
$200k – $300k per year
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on May 1, 2026. World Labs scores A on the Alion truth index.

Overview
Company
Impact
Profile match
World Labs is a spatial intelligence laboratory building large world models that generate and reason about three dimensional environments. Founded by the computer vision researcher Fei-Fei Li with collaborators from Stanford and Google, it argues that language alone cannot give machines an understanding of physical space. The company demonstrated systems that turn a single image into an explorable three dimensional scene and raised over a billion dollars in valuation within its first year; its team includes leading researchers in computer vision, graphics and generative modelling.

About World Labs

World Labs is a frontier AI research and product company advancing spatial intelligence, the next frontier beyond large language models. Co-founded by Dr. Fei-Fei Li, Justin Johnson and Ben Mildenhall, the company is pioneering world models that perceive, generate, reason, and interact with virtual and physical worlds.

The company’s flagship product, Marble, transforms text, images, and video into fully navigable 3D worlds, unlocking applications across gaming, film, architecture, robotics, and immersive digital experiences. Backed by leading investors and with over $1B raised, World Labs is assembling a world-class team at the intersection of AI research and real-world deployment.

Role Overview

We are looking for a Performance Engineer to make World Labs’ models train and serve as fast as the hardware allows.

Running large generative world models at scale is a novel systems problem. You will find the bottlenecks - in kernels, in the serving path, in the training loop, in how we use our GPUs - and eliminate them. Your ownership is technical and concrete: the throughput you unlock, the latency you cut, the utilization you win back, and the correctness you hold while doing it. You will work up and down the stack, from low-level tensor and kernel optimization to fleet-wide serving efficiency, in close partnership with the researchers whose models you are accelerating.

This is a hands-on, individual-contributor role. You will profile, design, build, and ship code directly.

What You Will Do

  • Optimize inference and serving end to end - latency, throughput, batching, caching, and scheduling - to serve our models efficiently at production scale.

  • Write and tune GPU kernels (CUDA, Triton) for hot paths; drive kernel fusion, memory- and bandwidth-bound optimization, and low-precision (FP8/INT8) execution.

  • Optimize training throughput and GPU utilization: parallelism strategies, communication/compute overlap, mixed precision, and eliminating pipeline stalls.

  • Build performance models, profiling workflows, and observability that make throughput, latency, cost, utilization, and their tradeoffs legible across the stack.

  • Own numerical correctness across precision, kernel, and hardware changes - treating correctness as part of performance, not separate from it.

  • Partner with researchers to productionize models for serving and to make experiments run faster and more reliably.

  • Where needed, work on the distributed systems that training and inference run on - but the core of the job is squeezing the most out of every GPU.

Key Qualifications

You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary.

  • Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation.

  • Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale.

  • Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.

  • Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.

  • Working knowledge of ML framework internals (PyTorch and/or JAX; torch.compile, XLA, or similar compiler paths).

  • Strong proficiency in Python, with the ability to drop into C++/CUDA (and Rust or Go) as the work demands.

  • High-ownership mindset - you measure yourself by throughput shipped and latency cut, not tickets closed.

Preferred Qualifications

  • Experience at an AI lab or ML-native company, optimizing systems used directly by researchers and productionizing research code.

  • Low-precision and numerics depth: FP8/INT8 quantization, mixed-precision, and detecting numerical regressions across hardware platforms.

  • Distributed systems for large-scale training and inference - collective communication (NCCL), interconnects (NVLink), model and tensor parallelism, and fault tolerance. A strong plus, but not a substitute for the core skills above.

  • Experience serving generative, diffusion, video, or 3D/spatial models - not just text LLMs.

  • Multi-accelerator experience (GPU plus TPU or Trainium) and partnering with hardware vendors on accelerator capabilities.

  • Building performance-modeling and observability frameworks for GPU utilization and cost.

Who You Are

  • Fearless Innovator: We need people who thrive on challenges and aren't afraid to tackle the impossible.

  • Resilient Builder: Impacting Large World Models isn't a sprint; it's a marathon with hurdles. We're looking for builders who can weather the storms of groundbreaking research and come out stronger.

  • Mission-Driven Mindset: Everything we do is in service of creating the best spatially intelligent AI systems, and using them to empower people.

  • Collaborative Spirit: We're building something bigger than any one person. We need team players who can harness the power of collective intelligence.

We're hiring the brightest minds from around the globe to bring diverse perspectives to our cutting-edge work. If you're ready to work on technology that will reshape how machines perceive and interact with the world - then World Labs is your launchpad.

Join us, and let's make history together.

Logistics:

Location: This role is based in San Francisco, California; we work in-person in the office 5 days a week.

Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

Benefits: We offer generous health, dental, vision and 401K benefits, flexible PTO, paid parental leave, and relocation support as needed.

Equal Employment Opportunity

World Labs is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law. We welcome all qualified applicants and are committed to providing reasonable accommodations throughout the hiring process upon request.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
721,262 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$13k – $33k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Tashkent
Python
Rust
C++
C++
PyTorch C++
AI/ML
LangChain
LlamaIndex
vLLM
Computer Vision
NLP
Ollama
Transformers
PyTorch
RAG
Mixture of Experts
Reranking
Time Series Forecasting
Hugging Face
EU AI Act
Machine Learning
DevOps
Docker
Nginx
Cybersecurity
GDPR
Apply
$15k – $40k per year (Estimated) • In office • Astana
Python
C++
Bash
C++
CMake
AI/ML
OpenCV
CUDA Toolkit
CUDA
DevOps
Git
Docker
Linux
Robotics
GStreamer
Apply
$33k – $81k per year (Estimated) • In office • 3+ years exp • Moscow
Python
C++
DevOps
NixOS
Debian
Docker
Ubuntu
QEMU
GitLab
Linux
Apply
$24k – $51k per year (Estimated) • In office • Full-Time • Saint Petersburg
Python
Verilog
C++
SystemVerilog
VHDL
MATLAB
MATLAB
Simulink
DevOps
CI/CD
Git
Chips/EDA
Siemens ModelSim
Xilinx Vivado
Synplify Pro
Apply
$96k – $231k per year (Estimated) • Remote • Full-Time • 4+ years exp • Bachelor's Degree • London
Python
SQL
DevOps
Rest API
Management
Google Workspace
Microsoft Office
Apply
Product Engineer 1 day ago
$275k – $325k per year • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
LoRA
Fine-tuning
PEFT
World Models
Apply
$250k – $325k per year • In office • Full-Time • 4+ years exp • San Francisco
Python
Rust
C++
AI/ML
CUDA Toolkit
Synthetic Data
CUDA
NCCL
NVLink
World Models
Machine Learning
Game Dev
Unreal Engine
Lumen
Nanite
Apply
$200k – $250k per year • In office • Full-Time • 8+ years exp • San Francisco
AI/ML
Claude
ChatGPT
Gemini
World Models
Management
Slack
Google Workspace
Apply
$250k – $350k per year • In office • Full-Time • 8+ years exp • San Francisco
AI/ML
Reinforcement Learning
Computer Vision
World Models
Physical AI
Robotics
Sim-to-Real
Imitation Learning
Reinforcement Learning
Chips/EDA
PoC Library
Apply
$250k – $350k per year • In office • Full-Time • 6+ years exp • PhD • San Francisco
Python
C++
C++
PyTorch C++
AI/ML
Reinforcement Learning
PyTorch
Vision-Language-Action
World Models
Machine Learning
Robotics
Sim-to-Real
Imitation Learning
Reinforcement Learning
Apply
$46k per year • In office • Full-Time • San Francisco
Management
Agile
Apply
$197k – $314k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco • Washington
AI/ML
Cursor
Claude Code
Function Calling
AI Agents
Langfuse
LangSmith
LLM
Braintrust
OpenAI Codex
Agentforce
Replit
LLM Evaluation
Tool Use
Management
Slack
Apply
$229k – $345k per year • Equity • Remote/Hybrid • Full-Time • 7+ years exp • San Francisco • Milpitas • Mountain View • Los Altos • Fremont
DevOps
BGP
OSPF
Apply
$139k – $235k per year • In office • Full-Time • San Francisco
Apply
$139k – $299k per year (Estimated) • Remote • Full-Time • San Francisco
AI/ML
Cursor
Claude Code
Time Series Forecasting
Physical AI
Machine Learning
Apply
See all jobs
This is one of many
721,262 more open roles from verified company boards, updated every day.