686,202open jobs
39,765companies
97,316added this week
Browse all
Salary
$17k – $32k per year (Estimated)
Location
In office (Beijing, Shanghai, Shenzhen)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

Join NVIDIA’s Cosmos Lab Infrastructure team to develop training and post-training systems for advanced Physical AI models, including world foundation models and robot policies. Our infrastructure connects training, inference, and evaluation with simulation and real-world robot interaction. You will work with a mentor on a focused project scoped to your experience and internship duration, implementing and evaluating systems improvements on real AI workloads using NVIDIA’s GPU infrastructure.

What you’ll be doing:

  • Develop and optimize training infrastructure for advanced Physical AI world models, supporting pre-training, supervised fine-tuning (SFT), and reinforcement learning (RL). Explore distributed parallelism, sharding, low-precision training, compute-communication overlap, and numerical consistency and efficient weight synchronization between training and inference.

  • Build Physical AI post-training and RL infrastructure supporting advanced training algorithms. Connect simulation or, where applicable, real-robot interaction with experience collection, rollout inference, reward computation, training, and evaluation. Optimize these workflows through partitioning, pipelining, data transfer, and synchronization across synchronous, asynchronous, or disaggregated execution.

  • Improve efficiency and scalability across training, inference, simulation, and evaluation through scheduling, placement, dynamic resource allocation, and load balancing, supporting heterogeneous resources, elasticity, and fault recovery.

  • Analyze and optimize system performance, working with researchers to investigate, support, and compare emerging Physical AI models, training workflows, and algorithms from a systems perspective. Use profiling, benchmarking, and performance modeling to identify bottlenecks and measure throughput, latency, GPU utilization, and policy freshness. Share findings through tested code, documentation, and technical presentations, and contribute to research publications where appropriate.

What we need to see:

  • Pursuing a Bachelor’s, Master’s, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field.

  • Strong Python and debugging skills, with systems fundamentals in concurrency, distributed execution, memory management, or data movement.

  • Practical experience in at least one area: training infrastructure, RL infrastructure, simulation or robotics integration, or inference infrastructure. Coursework, research, open-source projects, and internships all count.

  • Strong analytical and communication skills, curiosity, and a willingness to learn.

  • Experience in every listed area, prior access to large GPU clusters, and model architecture or learning algorithm research are not required.

Ways to stand out from the crowd:

  • Experience optimizing training infrastructure, including distributed parallelism, low-precision training, GPU memory efficiency, or compute-communication overlap.

  • Experience optimizing scheduling, placement, resource allocation, or data transfer across training, rollout, simulation, and evaluation.

  • Experience extending RL pipelines, integrating simulation environments or robot interfaces, or optimizing inference; GPU profiling, C++/CUDA development, and open-source contributions or research in ML systems are also valued.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
686,202 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Beijing
$76k – $188k per year • In office • Internship • PhD • Santa Clara
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
Reinforcement Learning
Multimodal AI
Diffusion Models
TensorFlow
PyTorch
CUDA
World Models
Embodied AI
Robotics
MuJoCo
Model Predictive Control
Imitation Learning
Reinforcement Learning
Apply
In office • Internship • Master's Degree
Python
Java
C++
DevOps
CI/CD
Git
Management
Agile
Apply
$108k – $178k per year • In office • Full-Time • Bachelor's Degree • Santa Clara • Seattle
Python
C++
AI/ML
Flash Attention
AI Agents
LLM
MLIR
Apache TVM
Apply
$76k – $188k per year • In office • Internship • PhD • Santa Clara • Westford • Seattle
C++
AI/ML
CUDA Toolkit
CUDA
Apply
$86k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
Python
PowerShell
DevOps
Splunk
VMWare
Azure
Windows Server
Nagios
Apply
$100k – $167k per year • In office • Full-Time • Bachelor's Degree • Santa Clara
Python
Apply
$108k – $178k per year • In office • Full-Time • Bachelor's Degree • Santa Clara • Seattle
Python
C++
AI/ML
Flash Attention
AI Agents
LLM
MLIR
Apache TVM
Apply
$76k – $188k per year • In office • Internship • PhD • Santa Clara • Westford • Seattle
C++
AI/ML
CUDA Toolkit
CUDA
Apply
$76k – $188k per year • In office • Internship • PhD • Santa Clara
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
Reinforcement Learning
Multimodal AI
Diffusion Models
TensorFlow
PyTorch
CUDA
World Models
Embodied AI
Robotics
MuJoCo
Model Predictive Control
Imitation Learning
Reinforcement Learning
Apply
$53k – $137k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
C++
AI/ML
InfiniBand
NVLink
DevOps
HPC
Apply
$20k – $46k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Beijing
Apply
In office • Beijing
Apply
$18k – $43k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Beijing
Apply
$18k – $53k per year (Estimated) • In office • Full-Time • Beijing
Apply
$34k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Shanghai • Beijing
Apply
See all jobs
This is one of many
686,202 more open roles from verified company boards, updated every day.