699,301open jobs
41,468companies
98,642added this week
Browse all
Salary
$37k – $92k per year (Estimated)
Location
In office (Shanghai)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

Joining NVIDIA's DGX Cloud Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on optimizing efficiency and resiliency of AI workloads, as well as developing scalable AI and Data infrastructure tools and services. Our objective is to deliver a stable, scalable environment for AI researchers, providing them with the necessary resources and scale to foster innovation. We are seeking an AI infrastructure software engineer to join our team. You'll be instrumental in designing, building, and maintaining AI infrastructure that enable large-scale AI training and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of AI systems.

As a senior DGX Cloud AI Infrastructure software engineer at NVIDIA, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now!

What you’ll be doing:

  • Develop infrastructure software and tools for large-scale AI, LLM, and GenAI infrastructure.

  • Develop and optimize tools to improve infrastructure efficiency and resiliency.

  • Root cause and analyze and triage failures from the application level to the hardware level

  • Enhance infrastructure and products underpinning NVIDIA's AI platforms.

  • Co-design and implement APIs for integration with NVIDIA's resiliency stacks.

  • Define meaningful and actionable reliability metrics to track and improve system and service reliability.

  • Skilled in problem-solving, root cause analysis, and optimization.

What we need to see:

  • Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.

  • Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).

  • Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.

  • Proven track record in building and scaling large-scale distributed systems.

  • Experience with AI training and inferencing and data infrastructure services.

  • Familiar in operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus, Loki).

  • Proficiency in programming languages such as Python, C/C++, script languages

  • Excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential.

Ways to stand out from the crowd:

  • Experience in working with the large scale AI cluster

  • Strong understanding of NVIDIA GPUs, network technologies (RDMA, IB, NCCL)

  • Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray

  • Experience and root cause analysis of failures and datacenter scale

  • Strong background in software design and development.

NVIDIA leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for exceptional people like you to help us accelerate the next wave of artificial intelligence.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
699,301 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$21k – $58k per year (Estimated) • In office • Moscow
Python
C++
DevOps
Ubuntu
Linux
Robotics
ROS2
Isaac Sim
Apply
In office • Internship • 2+ years exp • Des Moines
JavaScript
TypeScript
SQL
C#
C++
C#
.NET
Frontend
Angular
Management
Agile
Apply
$161k – $225k per year • Equity • Remote • 4+ years exp
Python
Kotlin
Databases
MySQL
DevOps
AWS
Kubernetes
Apply
$10k – $23k per year (Estimated) • In office • 1+ year exp • Nizhny Novgorod
Python
Databases
Apache Kafka
AI/ML
AI Agents
GigaChat
DevOps
CI/CD
Kubernetes
Management
Jira
QA
Pytest
Apply
Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Pune • Bengaluru
Python
DevOps
Splunk
Ansible
VLAN
BGP
OSPF
MPLS
Apply
$87k – $136k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • Poland • Switzerland • Germany • Netherlands • Ukraine
AI/ML
InfiniBand
NVLink
DevOps
HPC
BGP
OSPF
Apply
$44k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
AI Agents
CUDA
DevOps
Platform Engineering
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
$44k – $106k per year (Estimated) • In office • Full-Time • 2+ years exp • Master's Degree • Shanghai • Beijing
Python
AI/ML
VLM
PyTorch
LLM
NVIDIA NeMo
Machine Learning
DevOps
HPC
Apply
$42k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Beijing • Shanghai • Shenzhen
Python
C++
DevOps
Linux
Apply
$29k – $79k per year (Estimated) • In office • Full-Time • Shanghai
AI/ML
Physical AI
Apply
$44k – $110k per year (Estimated) • In office • Full-Time • 2+ years exp • Shanghai • Beijing
C
C++
C
Pthreads
C++
TensorFlow C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
TensorRT
OpenMP
TensorFlow
CUDA
cuDNN
MLIR
Apache TVM
CUTLASS
DevOps
CI/CD
Apply
$35k – $87k per year (Estimated) • In office • Full-Time • 5+ years exp • Shenzhen • Shanghai
C++
DevOps
Linux
Apply
$36k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Beijing • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
Reinforcement Learning
AI Agents
LLM
CUDA
Post-training
Physical AI
Machine Learning
Robotics
Reinforcement Learning
Apply
In office • Internship • PhD • Shanghai • Beijing • Shenzhen
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
Chips/EDA
PoC Library
Apply
See all jobs
This is one of many
699,301 more open roles from verified company boards, updated every day.