368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$200k – $400k per year
Location
In office (San Jose)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Figure AI is an American robotics company founded in 2022 that develops general-purpose humanoid robots for commercial and household work. Its Figure 02 and Figure 03 machines combine custom actuators, battery systems and hands with the in-house Helix vision-language-action model, which lets one neural network drive both perception and motor control from natural-language instructions. Headquartered in San Jose, California, the company runs the BotQ manufacturing facility for high-volume production and has piloted its robots on logistics and automotive assembly work with commercial partners.

Figure is an AI robotics company developing autonomous general-purpose humanoid robots. The goal of the company is to ship humanoid robots with human level intelligence. Its robots are engineered to perform a variety of tasks in the home and commercial markets. Figure is headquartered in San Jose, CA.

Figure's vision is to deploy autonomous humanoids at a global scale. Our Helix team is looking for an experienced AI Training Performance Engineer to take our model training to the next level. This role is focused on improving distributed training frameworks for large scale model training, optimizing GPU kernels, exploring the relative gains of different accelerator types and co-designing our models to maximize utilization of our hardware.

Responsibilities

  • Optimize training performance for a 100B+ parameter models across 100k+ GPUs.
  • Collaborate with the broader team on accelerator choice, cluster topology, scheduling, and hardware procurement decisions to inform future scaling.
  • Write and optimize custom kernels (Triton/CUDA)
  • Build tooling and dashboards for continuous performance monitoring, regression detection, and root-cause analysis across training jobs
  • Optimize data loading and preprocessing pipelines so I/O never gates the accelerators
  • Improve checkpointing, fault tolerance, and elastic restart so large jobs recover quickly from node failures without losing significant wall-clock time
  • Partner with researchers to co-design model architectures and training recipes that are performant at scale (e.g., activation checkpointing strategies, mixed precision, sequence packing)
  • Extend and contribute to kernel compilers (e.g., Triton, Gluon) to improve iteration speed and enable targeting of custom/non-NVIDIA accelerators
  • Build and extend agentic systems that automatically generate, benchmark, and iterate on custom kernels
  • Evaluate emerging accelerator architectures (AMD, TPU, SRAM-based ASICs, and other novel hardware) for fit with our training workloads, and lead proof-of-concept ports/benchmarks
  • Explore different model/data parallelisms (FSDP, context parallel, expert parallel, etc.) to determine optimal configuration per model size.

Requirements

  • Bachelor's or Master's degree in Computer Science, Computer/Electrical Engineering, or a related field
  • 3+ years in AI performance engineering, with significant time leading large-scale performance improvement projects
  • Deep understanding of GPU architecture and performance characteristics (memory bandwidth, compute-bound vs. memory-bound ops, occupancy)
  • Proficiency with profiling tools (Nsight Systems/Compute, PyTorch Profiler, HTA, or similar) and ability to translate traces into concrete optimizations
  • Solid grasp of collective communication (NCCL) and modern networking concepts (RDMA, NVLink, InfiniBand/RoCE, topology-aware placement).
  • Strong Python and CUDA/C++ skills; comfortable reading and modifying framework internals
  • Experience debugging performance regressions and instability at scale (stragglers, hangs, OOMs, numerical divergence)
  • Experience defining and reasoning about hardware-efficiency metrics (MFU/HFU) and using them to drive optimization priorities

Bonus Qualifications

  • Experience with heterogeneous or multi-datacenter training setups and cross-cluster orchestration
  • Contributions to open-source ML systems projects (PyTorch, Megatron-LM, vLLM, DeepSpeed, JAX, etc.)
  • Exposure to non-NVIDIA accelerators (AMD GPUs, TPU/Trainium/Inferentia, or custom silicon) and heterogeneous fleet management.

The US base salary range for this full-time position is between $200,000 - $400,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$71k – $112k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austria
C#
C++
Python
Visual Basic
Apply
$150k per year • In office • Full-Time • New York
C++
Python
Apply
$106k – $149k per year • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Overland Park
Python
SQL
Databases
Amazon Redshift
Analytics
Power BI
Tableau
Marketing
Salesforce
Apply
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$16k – $60k per year (Estimated) • In office • Full-Time • PhD • Mumbai
Python
SQL
Python
pySpark
Databases
Presto
Snowflake
AI/ML
Dagster
Prefect
Spark
DevOps
Amazon S3
AWS
CI/CD
Analytics
ETL/ELT
Power BI
Tableau
Apply
$80k – $90k per year • In office • Internship • San Jose
C++
Python
DevOps
CI/CD
RTOS
Robotics
EtherCAT
Apply
$150k – $350k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
$150k – $300k per year • In office • Full-Time • 4+ years exp • San Jose
C++
Python
AI/ML
Human-in-the-Loop
Robotics
GTSAM
Inverse Kinematics
Sensor Fusion
Teleoperation
Apply
$150k – $250k per year • In office • Full-Time • 8+ years exp • San Jose
MATLAB
Python
MATLAB
Simulink
Robotics
EtherCAT
Apply
User Support Lead 21 day ago
$84k – $238k per year (Estimated) • In office • 5+ years exp • San Jose
AI/ML
AI Agents
Marketing
Zendesk
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$147k – $265k per year (Estimated) • In office • Full-Time • Folsom • San Jose
Apply
$117k – $255k per year (Estimated) • In office • Full-Time • 8+ years exp • Richardson • Boise • Folsom • San Jose
Verilog
AI/ML
Claude
Apply
$124k – $208k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin • San Jose
C++
Python
SystemVerilog
Chips/EDA
Formal Verification
Apply
$116k – $253k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Boise • San Jose
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.