397,589open jobs
13,889companies
77,448added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara)
Seniority
Architect · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking a Compute Kernel Performance Architect who can develop, profile, and analyze CUDA workloads with a strong focus on GPU power behavior. In this role, you will create specialized workloads that exercise the GPU’s compute, memory, and I/O subsystems under demanding operating conditions. You will work closely with GPU architects, power architects, silicon validation engineers, and software teams to characterize workload behavior and influence the power architecture of future NVIDIA products. This position sits at the intersection of GPU architecture, high-performance software, and silicon characterization.

What You'll Be Doing:

  • Design and develop CUDA kernels and infrastructure that exercise worst-case power behavior across GPU compute, memory, and I/O subsystems.

  • Profile workloads to understand the relationship between kernel behavior, hardware utilization, performance, and power consumption.

  • Build workloads that generate controlled steady-state and transient power conditions across multiple GPU architectures.

  • Partner with GPU architects and silicon teams to identify functional units and workload patterns that require additional characterization.

  • Support power-stress methodology from pre-silicon modeling and simulation through post-silicon bring-up and validation.

What We Need to See:

  • MS or PhD or equivalent experience in Computer Science, Electrical Engineering, Computer Engineering, or a related field-or equivalent practical experience.

  • 5+ years of experience in CUDA programming, GPU kernel development, high-performance computing, or performance architecture.

  • Hands-on experience developing and optimizing GPU kernels, including work at the PTX or assembly level.

  • Experience with GPU performance-analysis tools such as Nsight Compute, Nsight Systems, nvprof, or equivalent tools.

  • Strong understanding of GPU build principles, including streaming multiprocessors, execution pipelines, memory hierarchy, synchronization, occupancy, and power states.

  • Excellent analytical, debugging, and communication skills.

  • Ability to work effectively across GPU architecture, software, silicon validation, and hardware engineering teams.

Ways to Stand Out from the Crowd:

  • Experience crafting GPU power-stress microbenchmarks or test-to-failure workloads.

  • Familiarity with Power Delivery Network concepts, including package and board-level behavior, impedance, inductance, decoupling, resonance, voltage droop, and overshoot.

  • Understanding of di/dt and how changes in current over time can compose voltage transients.

  • Experience with DVFS, AVFS, clock management, power states, or hardware noise-mitigation mechanisms.

  • Knowledge of how software workload patterns can interact with system-level power-delivery behavior.

Our team works at the core of NVIDIA’s GPU performance and power stack. We collaborate closely with Compute Architecture, Power Architecture, Silicon Solutions, circuit-design teams, and deep-learning software teams. The workloads, tools, and analysis produced by this team help validate current products and influence the build of upcoming NVIDIA GPUs.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 26, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
397,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$63k – $151k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
C++
Python
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
PyTorch
DevOps
HPC
Apply
$27k – $74k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Rust
AI/ML
CUDA
CUDA Toolkit
LLM
PyTorch
SGLang
TensorRT
Triton
vLLM
DevOps
HPC
Apply
$117k – $209k per year (Estimated) • Remote • Full-Time • 3+ years exp • High School Diploma • United States
JavaScript
Node JS
Python
SQL
Databases
BigQuery
Google BigQuery
AI/ML
Anthropic
CUDA
CUDA Toolkit
cuDNN
Embeddings
Google ADK
LLM Guardrails
NLP
OpenAI
Prompt Engineering
RAG
Vertex AI
Frontend
React.js
DevOps
CI/CD
GCP
Analytics
A/B Testing
Apply
Remote • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
Edge AI
ONNX
PyTorch
Quantization
ROCm
TensorRT
Triton
vLLM
Apply
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 3 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
$171k – $324k per year (Estimated) • Equity • In office • 15+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Cybersecurity
Zero Trust
Marketing
Instagram
LinkedIn
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
See all jobs
This is one of many
397,589 more open roles from verified company boards, updated every day.