698,623open jobs
40,860companies
106,276added this week
Browse all
Salary
$33k – $95k per year (Estimated)
Location
In office (Shanghai, Beijing, Shenzhen)
Seniority
Intern
Employment
Internship
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are looking for a motivated Deep Learning engineer to integrate advanced communication technologies into AI stacks like PyTorch, vLLM, SGLang, TRT-LLM, and veRL. You will be working with the team that developed communication libraries -- such as NCCL and NVSHMEM -- for scaling Deep Learning applications. Your customers will have diverse multi-GPU needs, ranging from training on scales up to 100K GPUs to inference at microsecond latency. Communication performance between GPUs directly affects AI applications. Your work in AI toolkits will simplify these challenges for the community. This is an excellent opportunity for someone with an AI background to push the state of the art in this field. Are you ready to contribute to innovative technologies and help realize NVIDIA's vision?

What you'll be doing:

  • Integrate new communication libraries features in AI frameworks: from PoC to performance analysis to production.

  • Perform deep analysis of AI workloads and frameworks to identify multi-GPU communication requirements and opportunities. Collaborate hands-on with teams working on the latest AI models.

  • Author custom communication or fused compute-communication kernels to showcase ultimate performance on NV platforms.

  • Conduct in-depth research to achieve SOL GPU performance.

  • Build fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.

  • Collaborate with a very dynamic team across multiple time zones.

What we need to see:

  • You are pursuing a M.S. or Ph.D. in CE/CS/EE with a strong background in communication, kernel authoring, and/or AI training/inference.

  • Rapid prototyping and development with Python, C++, CUDA or related DSLs (Triton, cuTe).

  • Solid understanding of LLM models and parallelisms.

  • Adaptability and passion to learn new areas and tools.

  • Flexibility to work and communicate effectively.

Ways to stand out from the crowd:

  • Development experience with frameworks such as PyTorch, JAX, TRT-LLM, vLLM, SGLang, or veRL.

  • Experience with DL communication patterns such as Expert Parallelism (EP), TP, DP & PP.

  • Experience with CUDA kernel optimization and profiling.

  • Experience with large-scale training or production inference stack.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
698,623 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$41k – $101k per year (Estimated) • In office • Full-Time • 2+ years exp • Shanghai • Beijing
C
C++
C
Pthreads
C++
TensorFlow C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
TensorRT
OpenMP
TensorFlow
CUDA
cuDNN
MLIR
Apache TVM
CUTLASS
DevOps
CI/CD
Apply
$25k – $70k per year (Estimated) • In office • Full-Time • Master's Degree • Shanghai
Python
C++
Perl
DevOps
HPC
Apply
$37k – $86k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
GCP
AWS
Docker
Kubernetes
Linux
Windows
Apply
$39k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
AI/ML
CUDA Toolkit
ONNX
OpenCL
TensorFlow
PyTorch
CUDA
Apply
$39k – $83k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bengaluru
Python
Go
Rust
DevOps
SLURM
CI/CD
ArgoCD
Jenkins
Kubernetes
Incident Management
Linux
DNS
DHCP
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Apply
$37k – $86k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
GCP
AWS
Docker
Kubernetes
Linux
Windows
Apply
$39k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
AI/ML
CUDA Toolkit
ONNX
OpenCL
TensorFlow
PyTorch
CUDA
Apply
$39k – $83k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bengaluru
Python
Go
Rust
DevOps
SLURM
CI/CD
ArgoCD
Jenkins
Kubernetes
Incident Management
Linux
DNS
DHCP
Apply
$41k – $101k per year (Estimated) • In office • Full-Time • 2+ years exp • Shanghai • Beijing
C
C++
C
Pthreads
C++
TensorFlow C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
TensorRT
OpenMP
TensorFlow
CUDA
cuDNN
MLIR
Apache TVM
CUTLASS
DevOps
CI/CD
Apply
$34k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
Python
TypeScript
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
RAG
LLM Guardrails
DevOps
Azure
CI/CD
Platform Engineering
Management
Agile
Apply
$33k – $76k per year (Estimated) • In office • Full-Time • Shanghai
AI/ML
Physical AI
Apply
$23k – $65k per year (Estimated) • In office • PhD • Shanghai
Management
Agile
Apply
$29k – $62k per year (Estimated) • In office • PhD • Shanghai
Management
Agile
Apply
$45k – $93k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Shanghai
Apply
See all jobs
This is one of many
698,623 more open roles from verified company boards, updated every day.