595,900open jobs
29,787companies
85,713added this week
Browse all
Salary
$200k – $300k per year
Location
In office (Mountain View)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match

At Rhoda AI, we’re building the next generation of generalist intelligent robots. We own the full robotics stack from high-performance hardware and robot systems to the infrastructure and state-of-the-art foundation world models that control our robots. Our robots are designed to be generalists capable of operating in complex, real-world environments and handling long-tail edge cases, made possible by our cutting edge research and end-to-end system design. We've raised over $450M and are investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up to make generalist robotics a reality.

About the Role

We're looking for a senior or staff-level Research Engineer or ML Systems Engineer to make our robot- learning pipeline reliable, reproducible, and measurable from end to end.

You will own the supported path from robot data collection through dataset generation, model training,

inference, and real-robot evaluation. You will build the validation systems, regression tests, observability,

and operating practices that allow researchers to trust their experiments and quickly distinguish model limitations from data, software, infrastructure, hardware, or evaluation failures.

This is a core research systems role - not generic DevOps, infrastructure support, or traditional QA. You will need to understand the semantics of robot data and model behavior, work directly with training and inference code, investigate failures on physical robots, and drive fixes across organizational boundaries.

We are hiring across senior and staff levels. Because this will be the first dedicated owner of the end-to-end

workflow, we are primarily looking for someone with staff-level ownership and systems judgment.

What You'll Do

  • Own the robot-learning pipeline end to end. Establish and maintain a trusted workflow spanning robot data

collection, data ingestion and compilation, post-training, checkpoint generation, inference, and real-robot evaluation.

  • Define the supported golden path. Maintain known-good combinations of code, datasets, configurations, checkpoints, robot software, hardware settings, task stations, and evaluation procedures.
  • Build automated validation at every interface. Develop checks for timestamp synchronization, sensor and action integrity, episode completeness, schema compatibility, dataset migrations, dataloader outputs, preprocessing behavior, model inputs, and configuration correctness.
  • Create end-to-end regression tests. Build representative smoke tests that exercise data compilation, training, checkpoint loading, inference, replay or simulation, and real-robot execution. Develop small-scale overfit and canary experiments that catch correctness regressions before expensive training runs begin.
  • Ensure training and inference consistency. Identify and prevent discrepancies in image processing, sensor normalization, temporal context, action representation, model configuration, and other transformations used across training and deployment.
  • Make failures observable and diagnosable. Build instrumentation and debugging tools that help determine whether a performance regression originated in data, model code, infrastructure, inference, robot software, hardware configuration, the physical environment, or evaluation execution.
  • Improve real-robot evaluation reliability. Partner with researchers and robot operations to establish stable benchmark stations, reference baselines, clear rubrics, repeatable trial protocols, operator procedures, and tracking of environmental variables that affect performance.
  • Lead cross-functional root-cause investigations. Drive ambiguous failures to resolution across Research, Data Infrastructure, Model Infrastructure, Software, and Robot Operations. Turn incidents and regressions into durable tests, monitors, documentation, and interface contracts.
  • Establish release and compatibility standards. Define the validation required before changes to robot software, data systems, training code, or inference systems become part of the supported research workflow.
  • Measure and improve research reliability. Track pipeline success rates, reproducibility, regression frequency, time to root cause, benchmark stability, and other metrics that reflect the health of the research workflow.

What We're Looking For

  • A strong track record owning, integrating, or debugging complex ML, robotics, autonomy, or other sensor-rich systems across multiple layers.
  • Excellent software engineering skills, including strong Python, maintainable system design, automated testing, CI, code review, and production-quality debugging practices.
  • Hands-on experience with PyTorch or a comparable ML framework. You should be able to train or fine-tune models, inspect datasets and batches, interpret losses and outputs, load checkpoints, and debug inference behavior.
  • Experience building or operating ML workflows that span data processing, training, evaluation, and deployment - not only one isolated part of the stack.
  • Strong systems-debugging ability. You can take an ambiguous report such as "robot performance dropped" and turn it into a structured investigation with controlled experiments, measurements, and a defensible root cause.
  • Familiarity with data-pipeline correctness, including schemas, versioning, temporal data, migrations, provenance, validation, and reproducibility.
  • Experience with Linux, containers, and modern compute environments such as Kubernetes, Slurm, or distributed GPU clusters.
  • Strong attention to detail and a willingness to investigate silent failures that do not produce obvious exceptions or crashes.
  • Ability to work effectively across research, infrastructure, software, and operations teams and to drive decisions without relying on formal managerial authority.
  • Clear written communication, including the ability to document interfaces, test plans, root-cause analyses, release criteria, and operational procedures.
  • Comfort working directly with physical systems and dealing with the variability and ambiguity inherent in real- world robotics.
  • A degree in Computer Science, Robotics, Electrical Engineering, or a related discipline - or equivalent practical experience. A PhD and publication record are not required.

Nice to Have, but Not Required

  • Experience in robotics, autonomous driving, drones, industrial automation, warehouse automation, or another domain that combines learned models with physical systems.
  • Familiarity with robot data collection, teleoperation, camera and sensor systems, calibration, proprioception, force/torque sensing, or action synchronization.
  • Experience with ROS or ROS2, real-time robot systems, simulation, replay infrastructure, or hardware-in-the- loop testing.
  • Experience with imitation learning, behavior cloning, robot post-training, reinforcement learning, video models, multimodal models, or learned control policies.
  • Experience building golden datasets, model-quality regression suites, data contracts, or automated dataset validation.
  • Familiarity with columnar datasets and data infrastructure, including formats such as Parquet and systems for indexing, querying, or migrating large datasets.
  • Experience debugging distributed training, GPU environments, checkpointing, networking, storage, Kubernetes, Slurm, or InfiniBand-related failures.
  • Experience designing statistically sound evaluation protocols for systems with noisy or stochastic performance.
  • Prior experience in ML platform engineering, research infrastructure, autonomy validation, release engineering, SRE, or systems integration.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
595,900 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Mountain View
$168k – $245k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • San Jose • Reston
Python
DevOps
Terraform
Ansible
Helm
GitOps
AWS
Kubernetes
Amazon EKS
Amazon S3
Apply
$33k – $74k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bengaluru • Noida • Chennai • Gurgaon • Hyderabad
Python
Go
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
Prometheus
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Google GKE
Azure AKS
Apply
$251k – $417k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
Python
Go
Java
C++
DevOps
Splunk
Kubernetes
Platform Engineering
Apply
$13k – $28k per year (Estimated) • Remote • Full-Time • 2+ years exp • Moscow
Python
Java
Bash
Java
Apache Tomcat
Databases
PostgreSQL
ScyllaDB
Apache Kafka
DevOps
Puppet
Ansible
Zabbix
Helm
Prometheus
HAProxy
CI/CD
GitOps
ArgoCD
Jenkins
Git
Docker
Kubernetes
Nginx
Grafana
GitLab
Management
Jira
Apply
$158k – $303k per year (Estimated) • In office • Full-Time • Master's Degree • Sunnyvale
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
JAX
Multimodal AI
Computer Vision
TensorRT
TensorFlow
PyTorch
Self-Supervised Learning
Human-in-the-Loop
Edge AI
Vision-Language-Action
Embodied AI
Robotics
Sim-to-Real
Apply
$175k – $250k per year • In office • Full-Time • 4+ years exp • Mountain View
AI/ML
Claude
World Models
Physical AI
Design
Figma
Apply
$175k – $250k per year • In office • Full-Time • 7+ years exp • Mountain View
AI/ML
World Models
Apply
$225k – $280k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Mountain View
Python
MATLAB
MATLAB
Simulink
AI/ML
World Models
Apply
$150k – $200k per year • In office • Full-Time • 3+ years exp • Mountain View
Python
C++
AI/ML
World Models
DevOps
Git
Apply
$175k – $250k per year • In office • Full-Time • Mountain View
AI/ML
World Models
Apply
$120k – $160k per year • Equity • Remote/Hybrid • Full-Time • 7+ years exp • Mountain View
AI/ML
NLP
Apply
$130k – $150k per year • In office • Full-Time • Bachelor's Degree • Mountain View • Los Angeles
Python
JavaScript
Kotlin
TypeScript
Dart
Databases
PostgreSQL
Redis
AI/ML
NLP
LLM
Frontend
React.js
Mobile
Flutter
Offline-First
DevOps
Kubernetes
Apply
$80k – $120k per year • In office • Internship • Bachelor's Degree • Mountain View
Python
JavaScript
Kotlin
TypeScript
Dart
Databases
PostgreSQL
Redis
AI/ML
NLP
Frontend
React.js
Mobile
Flutter
DevOps
Kubernetes
Apply
$133k – $338k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Atlanta • Milwaukee • Dallas • Columbus • Kirkland
Apply
Accounting Manager 1 day ago
$135k – $155k per year • In office • 6+ years exp • Bachelor's Degree • Mountain View
Apply
See all jobs
This is one of many
595,900 more open roles from verified company boards, updated every day.