410,350open jobs
14,253companies
73,765added this week
Browse all
Location
In office (Shanghai)
Seniority
Architect · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been redefining computer graphics, PC gaming, and accelerated computing for more than 25 years. Today, we lead in artificial intelligence, driving advances in natural language processing, computer vision, autonomous systems, and scientific research. We are looking for a forward-thinking HPC and AI Inference Software Architect to help shape the future of scalable AI infrastructure-focusing on distributed training, real-time inference, and communication optimization across large-scale systems.

Join our world-class team of researchers and engineers building next-generation software and hardware systems that power the most demanding AI workloads on the planet.

What you will be doing:

  • Design and prototype scalable software systems that optimize distributed AI training and inference-focusing on throughput, latency, and memory efficiency.

  • Develop and evaluate enhancements to communication libraries such as NCCL, UCX, and UCC, tailored to the unique demands of deep learning workloads.

  • Collaborate with AI framework teams (e.g., TensorFlow, PyTorch, JAX) to improve integration, performance, and reliability of communication backends.

  • Co-design hardware features (e.g., in GPUs, DPUs, or interconnects) that accelerate data movement and enable new capabilities for inference and model serving.

  • Contribute to the evolution of runtime systems, communication libraries, and AI-specific protocol layers.

  • Collaborate with customers to understand their needs and provide innovative solutions for them.

What we need to see:

  • Ph.D, Masters, or Bachelors in computer science, computer engineering, electrical engineering or a closely related field.

  • 5+ years of experience in DNNs, Scaling of DNNs, Parallelism of DNN frameworks, or deep learning training workloads.

  • Deep understanding of Inference and Training workloads and optimizations, like Prefill/Decode, data parallelism, Tensor parallelism, FDSP, etc...

  • Experience with AI network parallelism using collective libraries and RDMA/RoCE.

  • Background in algorithm design, system programming, and computer architecture.

  • Strong programming and software development skills.

  • Ability and flexibility to work and communicate effectively in a multi-national, multi-time-zone corporate environment.

Ways to stand out from the crowd:

  • Deep understanding of technology and passion for what you do.

  • Strong collaborative and interpersonal skills, specifically a proven ability to effectively guide and influence within a dynamic matrix environment.

  • Background with designing communication middleware for high-performance computing systems, including RoCE and DPUs.

  • Background with CUDA programming and NVIDIA GPUs and programming models for emerging architectures.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
410,350 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$60k – $300k per year • Equity 0.1–0.5% • Remote • Full-Time • 6+ years exp • San Francisco
C++
Elixir
JavaScript
Kotlin
Node JS
Python
Rust
Swift
TypeScript
AI/ML
CUDA
CUDA Toolkit
Edge AI
Embeddings
Semantic Search
Semantic Search
Frontend
npm
WebAssembly
DevOps
Vector
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
$43k – $111k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree
Python
AI/ML
Computer Vision
Apply
Senior AI Architect 2 days ago
$117k – $121k per year • In office • Full-Time • 4+ years exp • Master's Degree • San Francisco
Python
Databases
Databricks
AI/ML
Anthropic
Computer Vision
EU AI Act
LLMOps
OpenAI
DevOps
AWS
Azure
GCP
Terraform
Cybersecurity
GDPR
Apply
Senior AI Architect 2 days ago
$137k – $142k per year • In office • Full-Time • 4+ years exp • Master's Degree • San Francisco
Python
Databases
Databricks
AI/ML
Anthropic
Computer Vision
EU AI Act
LLMOps
OpenAI
DevOps
AWS
Azure
GCP
Terraform
Cybersecurity
GDPR
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
$144k – $230k per year • In office • Full-Time • PhD • Santa Clara
Apex
Apex
MuleSoft
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
In office • Full-Time • 5+ years exp • Master's Degree • Shanghai
C++
Python
AI/ML
AI Agents
Copilot
Cursor
LLM
Reinforcement Learning
Robotics
Isaac Lab
Isaac Sim
Perception
Reinforcement Learning
ROS
Sim-to-Real
Teleoperation
Apply
In office • PhD • Shanghai
Design
SolidWorks
Apply
Engineer (Ph.D.) 1 day ago
In office • PhD • Shanghai
Apply
In office • PhD • Shanghai
Apply
In office • PhD • Shanghai
Apply
In office • PhD • Shanghai
Apply
See all jobs
This is one of many
410,350 more open roles from verified company boards, updated every day.