697,753open jobs
40,891companies
106,255added this week
Browse all
Salary
$41k – $101k per year (Estimated)
Location
In office (Shanghai, Beijing)
Seniority
Senior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

The place to find available career opportunities at NVIDIA for you and people you know. We are now looking for a Senior Performance Software Engineer for Deep Learning Libraries! Do you enjoy tuning parallel algorithms and analyzing their performance? If so, we want to hear from you! As a deep learning library performance software engineer, you will be developing optimized code to accelerate linear algebra and deep learning operations on NVIDIA GPUs. The team delivers high-performance code to NVIDIA’scuDNN,cuBLAS, andTensorRT libraries to accelerate deep learning models. The team is proud to play an integral part in enabling the breakthroughs in domains such as image classification, speech recognition, and natural language processing. Join the team that is building the underlying software used across the world to power the revolution in artificial intelligence! We’re always striving for peak GPU efficiency on current and future-generation GPUs. To get a sense of the code we write, check out ourCUTLASS open-source project showcasing performant matrix multiply on NVIDIA’sTensor Cores with CUDA. This specific position primarily deals with code lower in the deep learning software stack, right down to the GPU HW.

What you'll be doing:

  • Writing highly tuned compute kernels to perform core deep learning operations (e.g. matrix multiplies, convolutions, normalizations)

  • Following general software engineering best practices including support for regression testing and CI/CD flows

  • Collaborating with teams across NVIDIA:

    • CUDA compiler team on generating optimal assembly code

    • Deep learning training and inference performance teams on which layers require optimization

    • Hardware and architecture teams on the programming model for new deep learning hardware features

What we need to see:

  • Masters or PhD degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or related field

  • 2+ years of relevant industry experience

  • Demonstrated strong C++ programming and software design skills, including debugging, performance analysis, and test design

  • Experience with performance-oriented parallel programming, even if it’s not on GPUs (e.g. with OpenMP or pthreads)

  • Solid understanding of computer architecture and some experience with assembly programming

  • Identify bottlenecks, optimize resource utilization, and improve throughput.

Ways to stand out from the crowd:

  • Tuning BLAS or deep learning library kernel code

  • CUDA GPU programming

  • Numerical methods and linear algebra

  • LLVM, TVM tensor expressions, or TensorFlow MLIR

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
697,753 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$33k – $95k per year (Estimated) • In office • Internship • PhD • Shanghai • Beijing • Shenzhen
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
Chips/EDA
PoC Library
Apply
$39k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
AI/ML
CUDA Toolkit
ONNX
OpenCL
TensorFlow
PyTorch
CUDA
Apply
$34k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Shenzhen • Shanghai
C++
DevOps
Linux
Apply
$34k – $84k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Beijing • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
Reinforcement Learning
AI Agents
LLM
CUDA
Post-training
Physical AI
Machine Learning
Robotics
Reinforcement Learning
Apply
$25k – $70k per year (Estimated) • In office • Full-Time • Master's Degree • Shanghai
Python
C++
Perl
DevOps
HPC
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Apply
$37k – $86k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
DevOps
GCP
AWS
Docker
Kubernetes
Linux
Windows
Apply
$39k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
AI/ML
CUDA Toolkit
ONNX
OpenCL
TensorFlow
PyTorch
CUDA
Apply
$39k – $83k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bengaluru
Python
Go
Rust
DevOps
SLURM
CI/CD
ArgoCD
Jenkins
Kubernetes
Incident Management
Linux
DNS
DHCP
Apply
$34k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Shenzhen • Shanghai
C++
DevOps
Linux
Apply
$34k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
Python
TypeScript
AI/ML
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
RAG
LLM Guardrails
DevOps
Azure
CI/CD
Platform Engineering
Management
Agile
Apply
$33k – $76k per year (Estimated) • In office • Full-Time • Shanghai
AI/ML
Physical AI
Apply
$23k – $65k per year (Estimated) • In office • PhD • Shanghai
Management
Agile
Apply
$29k – $62k per year (Estimated) • In office • PhD • Shanghai
Management
Agile
Apply
$45k – $93k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Shanghai
Apply
See all jobs
This is one of many
697,753 more open roles from verified company boards, updated every day.