368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office (Shanghai, Beijing)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

We are looking for a highly motivated software engineer intern for an exciting role in our communication libraries and network software team.The position will be part of a fast-paced crew that develops and maintains software for complex heterogeneous computing systems that power disruptive products in High Performance Computing and Deep Learning.

What you will be doing:

  • Design, implement and maintain highly-optimized communication runtimes for Deep Learning frameworks (e.g. NCCL for TensorFlow/Pytorch) andHPC programming interfaces (e.g. UCX for MPI/OpenSHMEM) on GPU clusters.

  • Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM.

  • Design, implement and maintain system software that enables interactions among GPUs and interactions between GPUs and other system components.

  • Creating proof-of-concepts to evaluate and motivate extensions in programming models, new designs in runtimes and new features in hardware.

What we need to see:

  • You are pursuing a Ph.D. in CE/CS/EE with a strong background in computer architecture, operating systems, communication library and/or AI/ML.

  • Excellent C/C++ programming and debugging skills.

  • Strong experience with Linux.

  • Experience with parallel programming interfaces and communication runtimes.

  • Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment.

Ways to stand out from the crowd:

  • Deep knowledge of high-performance networks like InfiniBand, RoCE etc.

  • Background with HPC applications. Experience with Deep Learning Frameworks such PyTorch, TensorFlow, JAX/XLA, vLLM/SGLang etc.

  • Experience with AI/DL communication patterns such as Expert Parallelism (EP), TP, DP, PP and how these patterns can be implemented with NCCL. Experience with CUDA kernel optimization and profiling.

  • Experience with large-scale model training and production inference software stack.

  • Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$140k – $253k per year (Estimated) • In office • 5+ years exp • Long Beach
C++
Rust
C++
Protobuf
AI/ML
Human-in-the-Loop
DevOps
CI/CD
Vector
Apply
$100k – $500k per year • In office • Full-Time • 1+ year exp • Fort Collins
C++
Python
SystemVerilog
Cython
C++
CMake
Cython
PyBind11
AI/ML
ChatGPT
Claude
Copilot
Edge AI
Chips/EDA
Synopsys ZeBu
Apply
$220k – $300k per year • In office • Full-Time • 3+ years exp • Mountain View
Python
AI/ML
Embeddings
JAX
MLFlow
PyTorch
TensorFlow
Weights & Biases
DevOps
AWS
Azure
GCP
Apply
$103k – $216k per year (Estimated) • In office • PhD • London
Python
AI/ML
Fine-tuning
JAX
LLM
PyTorch
LLM Guardrails
Cybersecurity
Least Privilege
Apply
$14k per year (net) • In office • Full-Time • Rostov-on-Don
C++
Python
C++
CMake
DevOps
GitHub
Apply
In office • Full-Time • 3+ years exp • Master's Degree • Shanghai • Shenzhen
C++
C
C++
TBB
C
MPI
Pthreads
AI/ML
CUDA
CUDA Toolkit
cuDF
OpenMP
RAPIDS
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
DevOps
AWS
Cybersecurity
GDPR
PCI DSS
Analytics
Power BI
Tableau
Management
Confluence
Jira
Trello
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
C++
Python
SQL
AI/ML
LangChain
LLM
Spark
Apply
In office • Full-Time • 15+ years exp • Bachelor's Degree • Shanghai
AI/ML
CUDA
CUDA Toolkit
LocalAI
DevOps
CI/CD
KVM
QEMU
Apply
In office • Full-Time • Shanghai
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.