459,924open jobs
15,541companies
67,347added this week
Browse all
Salary
$127k – $274k per year (Estimated)
Location
In office (Houston)
Employment
Full-Time
Overview
Company
Impact
Profile match
Persona AI is a humanoid robotics startup founded in 2024 to build industrial robots for shipbuilding and heavy fabrication. Its founding team came from Apptronik, NASA's Valkyrie programme and conversational artificial intelligence startups. The company works with Korean shipbuilder HD Hyundai on welding robots for dangerous shipyard tasks.

Job Title: AI Engineer - Robotics Data Preprocessing

Department: Software

Reports To: Teleoperations Lead

Employment Type: Full-Time

Location: Houston, TX or Pensacola FL

Who We Are

Persona AI is building humanoid robots for the most demanding environments in heavy industry - shipyards, steel mills, fabrication facilities, and offshore platforms - performing welding, grinding, maintenance, inspection, and material-handling work that is dangerous, physically demanding, and increasingly difficult to staff.

We are backed by leading strategic and financial investors and engaged with global industrial leaders across Korea, Japan, the United States, and Singapore. Korea is the center of gravity for our early commercial strategy, anchored by relationships with the world’s leading shipbuilders and steelmakers. Our work spans both the robot platform itself and the systems, partners, and playbooks required to deploy it at scale.

Why Join Persona AI?

  • We offer competitive compensation, a performance-based bonus, 99% employer covered medical benefits, early-stage equity, competitive PTO, and a company-wide paid winter break between December 24th and January 2nd.

  • You’ll shape technology that’s redefining the possibilities of robotics and human interaction.

  • Work alongside passionate teammates who value creativity, and continuous learning.

  • Enjoy full access to advanced tools,

About the Role

At Persona we require an unprecedented volume of high-quality, multimodal data. We are moving beyond basic teleoperation to leverage massive datasets of in-the-wild egocentric video combined with dense sensor streams (IMU, haptics, kinematics, and high-fidelity force profiles). We are seeking a highly skilled AI Engineer to architect the systems that turn this raw, unstructured multimodal data into high-fidelity training assets for our robots.

Models are only as good as the data they learn from. In humanoid robotics, that's not a platitude, it's the bottleneck. There's no Internet-scale corpus of robots manipulating the physical world. We have to create it. That's this role.

As an AI Data Engineer, you sit at the most leveraged point in our entire training pipeline: every model we ship is downstream of the data you build. If this role succeeds, our foundation models learn dexterity faster than anyone else's. If it fails, nothing else matters.

You will architect and scale the infrastructure that turns raw, messy reality into training-grade data, extracting, augmenting, and aligning human dexterous manipulation data from massive multi-sensor and egocentric video datasets. You'll build advanced pre-processing algorithms that recover what sensors can't directly see: quantifying grasp dynamics from force-torque signals, estimating contact forces from visual cues alone, reconstructing heavily occluded hand poses, and lifting 3D geometry out of 2D frames.

And because every minute of teleoperation data is expensive, you'll make each one count: using spatial, temporal, and cross-modal augmentation to multiply the value of everything our collection team captures. Your work directly determines how fast our models learn and how far they can go.

What You Will Be Doing

  • Force Analysis & Hidden State Inference: Design cross-modal validation systems that verify video, proprioception, force/haptic signals, and language annotations agree with each other, e.g., reprojecting robot state into the image plane to confirm video-state consistency, and VLM-assisted checks that instructions match observed behavior.

  • Kinematic Retargeting & Alignment: orchestrating hand-tracking, segmentation, depth estimation, 3D reconstruction, and pose-tracking modules; retargeting human demonstrations into robot trajectories; and running simulation-in-the-loop validation (kinematic feasibility, physics replay, motion-consistency filtering) so synthesized data is physically grounded, not just visually plausible.

  • Advanced Data Augmentation: Implement robust data augmentation strategies (spatial transformations, temporal scaling, synthetic viewpoints, and sensor noise injection) to expand expert trajectories and improve the robustness of our learning models.

  • Teleoperation Synchronization: unified state-action representations across differing embodiments, coordinate frames, rotation conventions, gripper/hand parameterizations, and sampling rates, with per-dimension validity masking and per-source normalization so that adding a new robot or sensor is a configuration change, not a rewrite.

  • Close the loop with data consumers: build the tooling that lets researchers query, visualize, and audit datasets (clip browsers, trajectory viewers, annotation review UIs), and turn model-failure analyses into new curation rules and targeted re-collection requests.

  • Multimodal Data Pipelines: Architect end-to-end ingestion pipelines that take raw, unstructured recordings (egocentric video, teleoperation sessions, third-party open datasets) and produce indexed, queryable, training-ready datasets. This includes temporal segmentation of long recordings into action clips, metadata and scene-graph extraction, embedding-based retrieval, and language annotation workflows.

What We Are Looking For

  • Education: M.S., or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field.

  • Programming & ML Frameworks: Deep expertise in Python and extensive experience with PyTorch, specifically in handling custom dataloaders for multimodal datasets.

  • Force & Time-Series Data Processing: Experience analyzing and processing complex time-series data from force-torque (F/T) sensors, load cells, or tactile arrays, ensuring pristine alignment with visual frames.

  • Video Processing Expertise: Mastery of video processing pipelines and libraries (OpenCV, FFmpeg, Decord) and managing the I/O bottlenecks of terabyte-scale video datasets.

  • Solid working knowledge of 3D geometry and robotics data: coordinate frames and transforms, rotation representations, camera intrinsics/extrinsics, forward/inverse kinematics, URDF.

  • Data Augmentation: Proven ability to implement programmatic and generative data augmentation techniques for computer vision and time-series data.

Bonus Skills

  • Experience with NVIDIA’s robotic software stack (Open X-Embodiment, DROID, AgiBot World, EgoDex, or similar).

  • Familiarity with the modern perception toolbox as a user: segmentation (SAM-family), monocular depth, hand/body pose estimation (MANO/SMPL), 6-DoF object pose tracking, point tracking-you don't need to train these models, but you should be comfortable composing and evaluating them in a pipeline

  • Familiarity with distributed data processing systems (Ray, Apache Spark) for cluster computing.

  • Background in generating or utilizing synthetic robotic data via simulation (Omniverse, MuJoCo).

  • Experience integrating spatial awareness or tactile data representations (e.g., Fourier encoding) into visual pipelines.

Persona AI is an Equal Opportunity Employer.

All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, age, disability, veteran status, or any other characteristic protected by applicable federal, state, or local law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
459,924 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Houston
$133k – $236k per year • Equity • In office • Full-Time • 7+ years exp • San Jose
Python
SQL
Databases
Weaviate
Databricks
Milvus
Pinecone
AI/ML
LangGraph
AutoGen
LangChain
Claude
Model Context Protocol
Chain-of-Thought
AI Agents
CrewAI
Gemini
LLM
RAG
OpenAI
Agentic Workflows
DevOps
Rest API
Azure
AWS
Docker
Kubernetes
Apply
$14k – $41k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Noida
Python
JavaScript
Java
TypeScript
SQL
Java
Maven
Gradle
AI/ML
Copilot
Cursor
Claude
ChatGPT
LLM
Frontend
React.js
Mobile
JUnit
DevOps
Rest API
GitHub Actions
CI/CD
Jenkins
Git
Docker
Shift-Left
GitHub
Cybersecurity
Shift-Left Security
QA
TestNG
Selenium
Playwright
Pytest
Apply
$37k – $68k per year (Estimated) • Equity • In office • Full-Time • Toulouse
Python
MATLAB
SAS
Apply
$129k – $253k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Portland • Chattanooga
Python
SQL
Analytics
Tableau
Power BI
Alteryx
Apply
$52k – $104k per year • In office • Confidential • Internship • 2+ years exp • Rochester
Python
DevOps
GitHub Actions
GitHub
QA
Robot Framework
Apply
$58k – $115k per year (Estimated) • In office • Full-Time • 4+ years exp • High School Diploma • Houston
Python
MATLAB
Apply
$84k – $235k per year (Estimated) • In office • Internship • Bachelor's Degree • Houston
Python
C++
AI/ML
Reinforcement Learning
Computer Vision
Robotics
ROS
Drake
MoveIt
Isaac Sim
MuJoCo
Motion Planning
Reinforcement Learning
Digital Twin
Apply
$62k – $126k per year (Estimated) • In office • Full-Time • 3+ years exp • Houston
Management
Google Workspace
Outlook
Apply
$79k – $167k per year (Estimated) • In office • Full-Time • 5+ years exp • Houston
Marketing
LinkedIn
Apply
$88k – $186k per year (Estimated) • In office • Full-Time • 5+ years exp • Houston
AI/ML
Reinforcement Learning
DevOps
GitHub
Robotics
Reinforcement Learning
Marketing
LinkedIn
Apply
$125k – $150k per year • In office • Full-Time • 7+ years exp • PhD • Amarillo • Charlotte • Las Vegas • Phoenix • Nashville
Apply
$94k – $237k per year (Estimated) • In office • Full-Time • 10+ years exp • Houston
AI/ML
Edge AI
Apply
$113k – $169k per year • Equity • Remote • Full-Time • 7+ years exp • High School Diploma • Jacksonville • Tempe • San Antonio • Dallas • Phoenix
Analytics
Power BI
Marketing
Salesforce
LinkedIn
Apply
$94k – $266k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Hartford • Milwaukee • Dallas • Columbus • Kirkland
Apply
$155k – $213k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree • New York • Houston
Python
TypeScript
AI/ML
Cursor
LangChain
Claude
Claude Code
LlamaIndex
LoRA
Fine-tuning
RLHF
Function Calling
AI Agents
NLP
PEFT
PyTorch
LLM
RAG
Hugging Face
OpenAI Codex
Tool Use
DevOps
GCP
Azure
AWS
Cybersecurity
HIPAA
Apply
See all jobs
This is one of many
459,924 more open roles from verified company boards, updated every day.