709,913open jobs
42,197companies
100,209added this week
Browse all
Salary
$17k – $33k per year (Estimated)
Location
In office (Shanghai, Beijing)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

The place to find available career opportunities at NVIDIA for you and people you know. We are now looking for a Performance Software Intern for Deep Learning Libraries! Do you enjoy tuning parallel algorithms and analyzing their performance? If so, we want to hear from you! As a deep learning library performance software intern, you will be developing optimized code to accelerate linear algebra and deep learning operations on NVIDIA GPUs. Join the team that is building the underlying software used across the world to power the revolution in artificial intelligence! We’re always striving for peak GPU efficiency on current and future-generation GPUs. To get a sense of the code we write, check out our CUTLASS open-source project showcasing performant matrix multiply on NVIDIA’s Tensor Cores with CUDA. This specific position primarily deals with code lower in the deep learning software stack, right down to the GPU HW.

What you'll be doing:

  • Writing highly tuned compute kernels to perform core deep learning operations (e.g. matrix multiplies, MoE, Attention)
  • Following general software engineering best practices including support for regression testing and CI/CD flows
  • Collaborating with teams across NVIDIA:

Compiler team on generating optimal assembly code

Deep learning training and inference performance teams on which layers require optimization

Hardware and architecture teams on the programming model for new deep learning hardware features

What we need to see:

  • Pursuing Masters or PhD degree in Computer Science, Computer Engineering, Applied Math, or related field
  • Demonstrated strong programming and software design skills, including debugging, performance analysis, and test design
  • Experience with performance-oriented parallel programming, even if it’s not on GPUs (e.g. with OpenMP or pthreads)
  • Solid understanding of computer architecture and some experience with assembly programming
  • Identify bottlenecks, optimize resource utilization, and improve throughput.

Ways to stand out from the crowd:

  • Tuning deep learning library kernel code
  • CUDA GPU programming
  • Numerical methods and linear algebra
  • LLVM, TVM tensor expressions, or TensorFlow MLIR
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
709,913 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$26k – $75k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Shanghai • Beijing • Shenzhen
Python
C++
AI/ML
CUDA Toolkit
CUDA
Machine Learning
DevOps
HPC
Apply
In office • Internship • Master's Degree • Beijing • Shanghai
Python
C++
C++
LLVM
AI/ML
CUDA Toolkit
CUDA
MLIR
Apply
$17k – $33k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
ChatGPT
TensorRT
TensorRT-LLM
TensorFlow
PyTorch
LLM
Machine Learning
Apply
$17k – $33k per year (Estimated) • In office • Internship • Master's Degree • Shanghai • Beijing
Python
C++
C++
PyTorch C++
LLVM
AI/ML
CUDA Toolkit
TensorRT
TensorRT-LLM
PyTorch
LLM
Mixture of Experts
CUDA
Edge AI
MLIR
Apply
$17k – $33k per year (Estimated) • In office • Internship • Master's Degree • Shanghai • Beijing
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
ChatGPT
TensorRT
TensorRT-LLM
TensorFlow
PyTorch
LLM
Machine Learning
Apply
$26k – $75k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Shanghai • Beijing • Shenzhen
Python
C++
AI/ML
CUDA Toolkit
CUDA
Machine Learning
DevOps
HPC
Apply
$60k – $129k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Taipei
Apply
$60k – $129k per year (Estimated) • In office • Full-Time • 5+ years exp • Hsinchu
Apply
$60k – $129k per year (Estimated) • In office • Full-Time • 5+ years exp • Hsinchu
Apply
In office • Internship • Master's Degree • Beijing • Shanghai
Python
C++
C++
LLVM
AI/ML
CUDA Toolkit
CUDA
MLIR
Apply
SAP FI Consultant 1 day ago
In office • Full-Time • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
Apply
In office • Full-Time • Guangzhou • Shenzhen • Shanghai • Dalian • Chengdu
Apply
$46k – $92k per year (Estimated) • In office • Full-Time • 10+ years exp • Master's Degree • Shanghai • Beijing
Verilog
VHDL
Apply
$22k – $60k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai • Guangzhou • Shenzhen • Dalian
Apply
$15k – $41k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Shanghai
Marketing
Salesforce
Apply
See all jobs
This is one of many
709,913 more open roles from verified company boards, updated every day.