Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs
Kanaria Tech builds KRM: a foundational AI system that powers mobile robots with social navigation, self-learning, and human-level awareness. We help robotics companies achieve high-level autonomy (L4–L5) through plug-and-play APIs that turn machines into intelligent, socially aware agents.

About Kanaria Tech

Kanaria Tech is a Tokyo-based frontier physical AI lab building the core AI capabilities for physically embodied intelligence. Our flagship technology is the Kanaria Robotic Model (KRM), an embodiment-agnostic, multimodal foundation model that gives robots socially aware and anticipatory navigation. We are starting with AMRs and will extend to other embodiments over time.

About the Physical AI Team

The Physical AI team owns KRM (the multimodal navigation foundation model) at the core of everything we ship. This is our frontier research and model-development group: it defines KRM's architecture, trains it at scale, and drives the roadmap toward social navigation.

The Role

We are looking for an AI engineer to help design, train, and advance KRM. You will work at the intersection of foundation models, computer vision, and robot learning. You'll build a model that perceives multimodal scenes, predicts how they will evolve, and produces both robot trajectories and human-readable interpretability outputs.

What You'll Do

  • Design, train, and iterate on KRM that ingests vision (plus LiDAR, radar, audio, and depth when available) and outputs navigation signal together with interpretability signals.
  • Develop and improve the observation encoders and the cross-attention world-state representation.
  • Improve the existing decoders that predict semantic segmentation, optical flow, depth, and camera pose for the current and future frames (the model's short-horizon world model).
  • Advance the camera-pose-prediction learning objective that helps to gather training data at scale.
  • Work on per-robot RL policies that adapt the shared model to specific hardware.
  • Drive roadmap items: better world representations, observation/action decoupling, language integration, and feedback loops.
  • Build and scale the large-scale video data pipeline that feeds training.
  • Build a benchmarking system for social robot navigation.

What We're Looking For

  • Strong foundation in deep learning and modern neural network architectures (transformers, attention).
  • Hands-on experience training large models in PyTorch (or an equivalent framework).
  • Solid Python engineering skills and comfort working with large-scale data.
  • Background in one or more of: computer vision, multimodal learning, robot learning, or world models.

Nice to Have

  • Experience with Vision-Language-Action (VLA) models, world models, video understanding, and 3D geometry / SLAM.
  • Reinforcement learning, especially sim-to-real or robot control.
  • Simulation experience (e.g., Isaac Sim) and synthetic data generation.
  • A research track record (publications) in relevant areas.

Recommended for you based on this role

Similar stack
Same company
In your city
Remote • Full-Time • Tokyo
C++
Python
DevOps
Docker
Apply
Remote • Full-Time • Tokyo
AI/ML
Multimodal AI
Robotics
ROS
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Tokyo
Remote work
Remote (Japan)
Employment
Full-Time