988,386open jobs
59,033companies
162,644added this week
Browse all
Location
In office (Shanghai)
Seniority
Intern
Employment
Internship

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 30, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is looking for outstanding Software Engineer Interns to help develop groundbreaking technologies for AI and deep learning kernel libraries. Our team builds core software that accelerates high-impact AI workloads on NVIDIA GPUs, with a strong focus on deep learning primitives, kernel libraries, and performance-critical GPU software. As an intern on the team, you will contribute to the design, development, optimization, and delivery of software that powers NVIDIA's AI platform.

This internship is centered on foundational library engineering, with opportunities to work on low-level kernels, performance primitives, and efficient implementations for modern AI and deep learning workloads. You may contribute to GPU-accelerated deep learning primitives, attention kernel implementations, runtime components, code generation systems, and other performance-critical infrastructure for large language models and advanced AI applications. You will collaborate with world-class engineers across deep learning software, compilers, GPU architecture, and open-source inference ecosystems, and your work can directly impact the performance of real-world workloads at scale.

What you'll be doing

  • Contribute to production-quality software that ships as part of NVIDIA's AI software stack, including cuDNN, FlashInfer, and optimized support for large language model inference workloads.
  • Help develop new AI systems technologies for efficient inference, with a focus on performance, scalability, maintainability, and usability.
  • Support the design, implementation, and optimization of kernels for high-impact AI workloads across LLM inference, generative AI, computer vision, autonomous driving, and recommender systems.
  • Assist in building extensible software abstractions for deep learning libraries, LLM serving engines, and runtime systems.
  • Contribute to just-in-time compilation, code generation, and runtime technologies for performance-critical GPU workloads.
  • Analyze workload performance, tune current software, and help propose improvements to future software and hardware-software interfaces.
  • Collaborate closely with engineers across deep learning frameworks, libraries, kernels, compilers, and GPU architecture teams at NVIDIA.
  • Contribute to open-source communities and ecosystem integrations where relevant, including projects such as FlashInfer, vLLM, and SGLang.

What we need to see

  • Currently pursuing a Bachelor's, Master's, or PhD degree in Computer Science, Electrical Engineering, or a related field.
  • Coursework, research, or hands-on project experience in machine learning, deep learning systems, compilers, systems software, or GPU programming.
  • Strong programming skills in C/C++ and Python.
  • Familiarity with CUDA development and GPU programming fundamentals.
  • Experience developing with or using deep learning frameworks such as PyTorch, JAX, TensorFlow, or ONNX.
  • Understanding of linear algebra, performance analysis, profiling, and code optimization.
  • Interest in software abstractions, APIs, and higher-level system architecture for performance-sensitive systems.
  • Interest in modern machine learning and inference system trends, especially around LLMs and generative AI.
  • Strong problem-solving skills, curiosity, and the ability to work effectively in a collaborative environment.

Ways to stand out from the crowd

  • Hands-on experience with inference engines and runtimes such as vLLM, SGLang, MLC, TensorRT-LLM, or similar systems.
  • Background in domain-specific compilers, code generation, or library solutions for LLM inference and training.
  • Exposure to machine learning compilers or IR systems such as MLIR, Apache TVM, TensorIR, or related technologies.
  • Practical experience with GPU performance modeling, computer architecture, or accelerator-oriented software design.
  • Open-source project ownership or meaningful contributions in deep learning systems, compilers, kernels, or inference infrastructure.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
988,386 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Shanghai
In office • Full-Time • 3+ years exp • Beijing
Python
Go
C++
Databases
Redis
ClickHouse
LanceDB
AI/ML
Ray
DevOps
Git
Apply
Agent 架构师 8 hours ago
≈ $25k – $68k per year (Estimated) • In office • Full-Time • 8+ years exp • Beijing
Apply
全栈开发工程师 8 hours ago
In office • Full-Time • Beijing
Python
TypeScript
AI/ML
LLM
Apply
In office • Full-Time • Beijing
Python
JavaScript
Java
Node JS
AI/ML
Function Calling
LLM
Frontend
Vue.js
React.js
Apply
≈ $25k – $68k per year (Estimated) • In office • Full-Time • 8+ years exp • Beijing
Apply
$64k per year • Hybrid • Internship • Bachelor's Degree • Chicago
Python
JavaScript
SQL
C++
DevOps
Linux
Management
Agile
Apply
$64k per year • Hybrid • Internship • Bachelor's Degree • Irving
Python
JavaScript
SQL
C++
DevOps
Linux
Management
Agile
Apply
$140k – $160k per year • Equity 0.1–0.2% • In office • Full-Time • 1+ year exp • New York
Python
JavaScript
SQL
Databases
PostgreSQL
AI/ML
AI Agents
Frontend
Next.js
React.js
DevOps
AWS
Marketing
LinkedIn
Apply
≈ $19k – $45k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Bengaluru
Python
SQL
Apply
≈ $29k – $75k per year (Estimated) • Remote (Colombia, Mexico, Costa Rica) • 5+ years exp • Mexico City
Python
PowerShell
DevOps
Terraform
Azure DevOps
Dynatrace
Azure
CI/CD
Bicep
Incident Management
SLI/SLO/SLA
Cybersecurity
Microsoft Entra ID
Active Directory
Management
ServiceNow
Agile
ITIL
Apply
In office • Internship • Master's Degree • Beijing • Shanghai
Python
C++
C++
LLVM
AI/ML
CUDA Toolkit
CUDA
MLIR
Apply
In office • Internship • PhD • Shanghai • Beijing • Shenzhen
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
Chips/EDA
PoC Library
Apply
Hybrid • Full-Time • Bachelor's Degree • Yokneam
Python
C
C++
C
FFmpeg
C++
CMake
AI/ML
CUDA Toolkit
CUDA
DevOps
Git
Linux
Robotics
GStreamer
Apply
≈ $83k – $137k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Warsaw
Python
Bash
AI/ML
InfiniBand
DevOps
Kibana
GitLab CI
CI/CD
Jenkins
Docker
Grafana
KVM
Linux
Apply
≈ $108k – $195k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Munich
C++
AI/ML
Computer Vision
DevOps
Linux
Apply
Production Manager 1 day ago
In office • Full-Time • Bachelor's Degree • Shanghai
Analytics
Microsoft Excel
Apply
≈ $52k – $111k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Shanghai
Apply
≈ $50k – $109k per year (Estimated) • In office • Full-Time • 10+ years exp • Master's Degree • Shanghai
Python
JavaScript
AI/ML
Model Context Protocol
Multimodal AI
AI Agents
TensorFlow
PyTorch
RAG
Machine Learning
Management
Agile
Apply
≈ $19k – $45k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Shanghai
Apply
≈ $18k – $43k per year (Estimated) • In office • Full-Time • 4+ years exp • Shanghai
Analytics
Master Data Management
Management
Agile
Apply
See all jobs
This is one of many
988,386 more open roles from verified company boards, updated every day.