1,161,622open jobs
13,727companies
227,717added this week
Browse all
Salary
$152k – $242k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars.

We are the GPU Communications Libraries and Networking team at NVIDIA. We deliver libraries like NCCL, NVSHMEM, UCX for Deep Learning and HPC. We are looking for a motivated Performance engineer to influence the roadmap of our communication libraries. The DL and HPC applications of today have a huge compute demand and run on scales which go up to tens of thousands of GPUs. The GPUs are connected with high-speed interconnects (eg. NVLink, PCIe) within a node and with high-speed networking (eg. Infiniband, Ethernet) across the nodes. Communication performance between the GPUs has a direct impact on the end-to-end application performance; and the stakes are even higher at huge scales! This is an outstanding opportunity for someone with HPC and performance background to advance the state of the art in this space. Are you ready for to contribute to the development of innovative technologies and help realize NVIDIA's vision?

What you will be doing:

  • Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.

  • Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack

  • Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available

  • Triage and root-cause performance issues reported by our customers

  • Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information

  • Collaborate with a very dynamic team across multiple time zones

What we need to see:

  • M.S. (or equivalent experience) or PhD in Computer Science, or related field with relevant performance engineering and HPC experience

  • 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)

  • Experience conducting performance benchmarking and triage on large scale HPC clusters

  • Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)

  • Implement micro-benchmarks in C/C++, read and modify the code base when required

  • Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python

  • Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker)

  • Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones

Ways to stand out from the crowd:

  • Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, congestion control

  • Experience debugging network issues in large scale deployments

  • Familiarity with CUDA programming and/or GPUs

  • Experience with Deep Learning Frameworks such PyTorch, TensorFlow

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 4, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,161,622 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$107k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
SQL
Databases
MySQL
DevOps
Ansible
AppDynamics
AWS
Blue-Green Deployment
CI/CD
Datadog
Docker
Dynatrace
Grafana
Jenkins
Kubernetes
New Relic
Prometheus
Terraform
Robotics
Digital Twin
Apply
$63k – $151k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
C++
Python
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
PyTorch
DevOps
HPC
Apply
AI Lead 3 days ago
$35k – $83k per year (Estimated) • In office • Bengaluru
Python
Python
FastAPI
Flask
Databases
Chroma
Databricks
FAISS
Pinecone
Weaviate
AI/ML
AI Agents
Amazon SageMaker
AutoGen
AWS Bedrock
Claude
CrewAI
Embeddings
Fine-tuning
Function Calling
Gemini
Hugging Face
Kubeflow
LangChain
LangGraph
LlamaIndex
LLM
LLMOps
LoRA
MLFlow
NLP
OpenAI
PEFT
Prompt Engineering
PyTorch
RAG
Semantic Kernel
Semantic Search
Semantic Search
TensorFlow
Transformers
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Rest API
Vector
Analytics
ETL/ELT
Apply
$88k – $168k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
SQL
TypeScript
Databases
PostgreSQL
DevOps
Docker
GCP
Kubernetes
Rest API
Apply
$55k – $125k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
SQL
TypeScript
Databases
PostgreSQL
DevOps
Docker
GCP
Kubernetes
Rest API
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
$107k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
SQL
Databases
MySQL
DevOps
Ansible
AppDynamics
AWS
Blue-Green Deployment
CI/CD
Datadog
Docker
Dynatrace
Grafana
Jenkins
Kubernetes
New Relic
Prometheus
Terraform
Robotics
Digital Twin
Apply
$63k – $151k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
C++
Python
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
PyTorch
DevOps
HPC
Apply
$171k – $324k per year (Estimated) • Equity • In office • 15+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Cybersecurity
Zero Trust
Marketing
Instagram
LinkedIn
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
See all jobs
This is one of many
1,161,622 more open roles from verified company boards, updated every day.