818,177open jobs
52,655companies
132,591added this week
Browse all
Salary
≈ $136k – $263k per year (Estimated)
Location
In office (Santa Clara)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
Powering generalist robotics and world models with multi-million-hour annotated egocentric data, hand tracking, upper-body/whole-body motion capture, tactile sensing, and simulation for OpenAI, Google DeepMind, Meta, Figure, 1X, Skild, Genesis AI, and Dyna Robotics.

Job Description:

We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale - turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Responsibilities

  • Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.

  • Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

  • Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.

  • Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.

  • Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.

  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.

  • Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.

  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.

  • Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.

  • Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.

Minimum Qualifications

  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.

  • 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.

  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.

  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.

  • Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).

  • Familiarity with agentic system design - tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).

  • Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).

Preferred Qualifications

  • PhD with a research focus in video understanding, multi-modal learning, or vision-language models.

  • Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.

  • Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.

  • Track record of deploying production-scale video search or retrieval systems.

  • Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.

  • Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).

  • Experience with humanoid robotics or embodied AI data pipelines is a plus.

Default Benefits:

  • Health insurance

  • Vision care

  • Dental coverage

  • 401(k)

  • Paid holidays

  • PTO (Paid Time Off)

  • Sick leave

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
818,177 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Santa Clara
≈ $138k – $268k per year (Estimated) • In office • Full-Time • 4+ years exp • PhD • San Jose
Python
Java
Ruby
C++
Scala
MATLAB
AI/ML
Machine Learning
Apply
≈ $109k – $247k per year (Estimated) • In office • Full-Time • 2+ years exp • San Jose
Python
C++
C++
TensorFlow C++
AI/ML
XGBoost
TensorFlow
Machine Learning
Apply
≈ $137k – $266k per year (Estimated) • In office • Full-Time • 3+ years exp • Seattle
Python
C++
C++
TensorFlow C++
AI/ML
XGBoost
TensorFlow
Machine Learning
Apply
≈ $152k – $276k per year (Estimated) • In office • Full-Time • Seattle
Python
C++
C++
TensorFlow C++
AI/ML
TensorFlow
Recommender Systems
Machine Learning
Apply
≈ $158k – $286k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Jose
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Computer Vision
TensorFlow
PyTorch
Machine Learning
Apply
Data Engineer 1 day ago
≈ $17k – $45k per year (Estimated) • In office • Full-Time • Kuala Lumpur
Python
Java
SQL
Databases
Snowflake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Hadoop
Spark
Airflow
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Analytics
ETL/ELT
Informatica
Talend
Azure Data Factory
Dimensional Modeling
Apply
$132k – $185k per year • Equity 1–3% • Hybrid • Full-Time • 6+ years exp • London
TypeScript
AI/ML
Cursor
LangChain
DSPy
Fine-tuning
AI Agents
LLM
OpenAI
Anthropic
Mastra
DevOps
Rest API
Management
Airtable
Apply
≈ $24k – $59k per year (Estimated) • In office • São Paulo
Python
SQL
Analytics
Power BI
Looker
Management
Google Sheets
Apply
≈ $46k – $110k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Chiyoda
Python
TypeScript
Python
FastAPI
AI/ML
AI Agents
LLM
RAG
Semantic Search
LLMOps
Semantic Search
Recommender Systems
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Apply
$117k – $177k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Washington
Python
Java
AI/ML
Copilot
Cursor
Claude Code
Prompt Engineering
AI Agents
LLM
OpenAI Codex
Agentforce
DevOps
Rest API
gRPC
Terraform
Helm
AWS
Docker
Kubernetes
IAM
Cybersecurity
HashiCorp Vault
Open Policy Agent
OWASP Top 10
SOC 2
Least Privilege
Cryptography
Vault
Apply
AI Agent Architect 1 month ago
≈ $176k – $359k per year (Estimated) • In office • Full-Time • Santa Clara
AI/ML
Multimodal AI
Function Calling
AI Agents
LLM
Human-in-the-Loop
LLM Guardrails
Tool Use
Apply
≈ $141k – $274k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
Multimodal AI
Computer Vision
TensorFlow
PyTorch
Machine Learning
Robotics
Localization
Inverse Kinematics
Teleoperation
Imitation Learning
Apply
≈ $65k – $142k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Santa Clara
AI/ML
Physical AI
Embodied AI
Apply
≈ $89k – $199k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
AI/ML
Machine Learning
Management
Google Sheets
Apply
Key Account Executive 22 days ago
≈ $94k – $210k per year (Estimated) • In office • Full-Time • 3+ years exp • Santa Clara
AI/ML
Physical AI
Apply
$167k – $251k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
AI/ML
Function Calling
AI Agents
Post-training
Tool Use
World Models
Machine Learning
Apply
$110k – $120k per year • Hybrid • Contractor • Santa Clara
JavaScript
AI/ML
Model Context Protocol
Hugging Face
Frontend
npm
DevOps
Linux
Windows
Apply
$150k – $160k per year • Hybrid • 6+ years exp • Bachelor's Degree • Santa Clara
SQL
Databases
Google BigQuery
BigQuery
Management
Google Sheets
Apply
$86k – $96k per year • Hybrid • 5+ years exp • Santa Clara
Apply
$170k – $180k per year • Hybrid • Contractor • 4+ years exp • Bachelor's Degree • Santa Clara
Apply
See all jobs
This is one of many
818,177 more open roles from verified company boards, updated every day.