702,799open jobs
41,484companies
101,606added this week
Browse all
Salary
$200k – $300k per year
Location
In office (Mountain View)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Rhoda AI is a robotics company developing generalist intelligent robots, owning the full stack from high-performance robot hardware and systems to the infrastructure and video-based foundation world models that control them. Its Direct Video Action approach pre-trains models on over a million web videos and then post-trains on robot action data for industrial tasks such as logistics returns processing; the company has raised over $450 million and set up Rhoda AI Europe GmbH as its first international entity, with first field deployments planned for late 2026. It hires research staff, machine learning and inference engineers, and robotics, firmware and mechanical engineers.

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

We're looking for a Staff / Principal ML Training Systems Engineer to own training systems performance end-to-end. You will define how our models train at scale - driving efficiency, scalability, and correctness across large-scale multimodal training. This is a core systems role, not infrastructure support. Your work directly determines how efficiently we use compute, how well models scale across thousands of GPUs, and how quickly research can iterate.

What You'll Do

Own training performance end-to-end

  • Diagnose and improve performance of large-scale multimodal training (vision, video, proprioception, actions, language)

  • Build systematic performance attribution: step-time decomposition (compute vs communication vs input pipeline), scaling curves across cluster sizes, and bottleneck identification and prioritization

  • Drive measurable gains in:

    • Distributed efficiency (comm/compute overlap, bucketization, topology-aware mapping, parallelism strategies)

    • Compute efficiency (kernel hotspots, operator fusion, attention optimization, framework/runtime overhead)

    • Memory efficiency (activation checkpointing, sequence packing/bucketing, fragmentation reduction)

Design training systems (not just tune them)

  • Define and evolve parallelism strategies: data / tensor / pipeline / sharding / hybrid approaches

  • Improve execution efficiency through communication scheduling and overlap, graph capture and execution optimization, and runtime-level improvements

  • Contribute to and extend training frameworks where needed

Make performance observable and measurable

  • Establish source-of-truth performance metrics: step-time breakdowns, MFU / throughput / scaling efficiency

  • Build tools to identify bottlenecks quickly, track performance across model families, and compare scaling behavior across configurations

  • Develop regression detection: microbenchmarks, performance baselines, and automated detection of efficiency regressions

Partner deeply with researchers

  • Work side-by-side with research scientists and research engineers - no silos

  • Translate model innovations into scalable, efficient implementations

  • Advise on training tradeoffs for robotics world models: long-horizon sequences, rollout/evaluation cadence, multimodal and variable-length data

Collaborate on cluster-level efficiency

  • Work with infrastructure/SRE teams to improve utilization across large distributed jobs, impact of network and collective performance on training, and topology-aware job placement and scaling behavior

What We're Looking For

  • Proven track record improving large-scale distributed training performance

  • Deep hands-on experience with modern ML stacks (PyTorch required; JAX a plus)

  • Strong understanding of data / tensor / pipeline parallelism, sharded training (FSDP / ZeRO-style), communication patterns and overlap strategies, and scaling behavior across large GPU clusters

  • Strong systems intuition - ability to reason across compute, communication, and memory bottlenecks

  • Exceptional debugging and measurement ability: turn "training is slow" into clear bottlenecks, experiments, and validated improvements

  • High ownership mindset and comfort in a fast-moving environment

Nice to Have (But Not Required)

  • GPU kernel or compiler-level experience (CUDA, Triton, graph capture, operator fusion)

  • Experience with multimodal or video training (variable-length sequences, packing/bucketing)

  • Experience working on large-scale training frameworks or distributed runtimes

  • Familiarity with cluster topology, networking, and large-scale scheduling effects

Why This Role

  • Direct leverage on research velocity - every efficiency gain you make accelerates model iteration across the entire research team

  • Own the scalability and performance of large-scale multimodal training for real-world embodied intelligence, not static benchmarks

  • Improvements you make compound across every training run the company executes - high ownership, high impact, small elite team

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
702,799 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Mountain View
$136k – $307k per year (Estimated) • Equity • In office • 6+ years exp • Bachelor's Degree • Bellevue
Python
C++
C++
TensorFlow C++
PyTorch C++
Databases
Qdrant
ElasticSearch
AI/ML
Kubeflow
TensorFlow
PyTorch
Ray
Triton
Amazon SageMaker
Feature Store
DevOps
AWS
Management
Agile
Apply
$106k – $148k per year • Remote • 15+ years exp
Python
Databases
PostgreSQL
Snowflake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Cursor
Claude
Spark
MLFlow
Vertex AI
dbt
Prefect
TensorFlow
PyTorch
LLM
Hugging Face
Amazon SageMaker
Feature Store
DevOps
Terraform
GCP
GitHub Actions
CloudFormation
Pulumi
Azure
CI/CD
AWS
Docker
Kubernetes
AWS Lambda
Amazon EventBridge
Cybersecurity
SOC 2
GDPR
HIPAA
Analytics
ETL/ELT
Apply
LLM Model Developer 2 days ago
$44k – $111k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru • Mumbai • Chennai • Hyderabad • Pune
Python
AI/ML
Fine-tuning
TensorFlow
PyTorch
LLM
Apply
$265k – $335k per year • Remote/Hybrid • 8+ years exp • Bachelor's Degree • San Francisco
AI/ML
Multimodal AI
Anthropic
Interpretability
Apply
In office • Bachelor's Degree
Python
SQL
AI/ML
Scikit-learn
NLP
TensorFlow
Pandas
NumPy
PyTorch
Machine Learning
Analytics
Power BI
QlikSense
Informatica
Apply
$175k – $250k per year • In office • Full-Time • 5+ years exp • Mountain View
Python
AI/ML
World Models
DevOps
Docker
SLI/SLO/SLA
Linux
Robotics
ROS
ROS2
Apply
Senior UX/UI Designer 11 days ago
$175k – $250k per year • In office • Full-Time • 4+ years exp • Mountain View
AI/ML
Claude
World Models
Physical AI
Design
Figma
Apply
$225k – $280k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Mountain View
Python
MATLAB
MATLAB
Simulink
AI/ML
World Models
Apply
$150k – $200k per year • In office • Full-Time • 3+ years exp • Mountain View
Python
C++
AI/ML
World Models
DevOps
Git
Linux
Apply
$175k – $250k per year • In office • Full-Time • Mountain View
AI/ML
World Models
Apply
Community Manager 1 day ago
$56k – $101k per year (Estimated) • In office • 2+ years exp • High School Diploma • Mountain View
Apply
GTM Intern 2 days ago
$41k – $83k per year (Estimated) • Remote • Internship • Mountain View
Apply
$200k – $280k per year • Equity 0–0.2% • Remote/Hybrid • Full-Time • 3+ years exp • Mountain View
AI/ML
AI Agents
Apply
$132k – $227k per year • In office • 8+ years exp • Bachelor's Degree • Mountain View
DevOps
AWS
Apply
$209k – $360k per year • In office • 12+ years exp • Bachelor's Degree • Mountain View
Python
Java
DevOps
Azure
IAM
DNS
Cybersecurity
Okta
Zero Trust
Microsoft Entra ID
Active Directory
LDAP
Apply
See all jobs
This is one of many
702,799 more open roles from verified company boards, updated every day.