1,174,835open jobs
13,798companies
228,726added this week
Browse all
Salary
$124k – $196k per year
Location
In office (Santa Clara)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

The AI revolution is not powered by models alone, rather it advances when enormous amounts of computation become fast, efficient, and economical enough to turn new ideas into products people can use on a global scale. Faster training lets research and product teams test the next idea sooner. Lower-latency, higher-throughput inference makes AI assistants and agents more responsive and practical for more people. Shorter time to solution lets scientists and engineers explore more possibilities within the same time and energy budget.

At NVIDIA, performance is not a supporting metric - it is how architectural invention becomes useful computing. CUDA is a critical layer where that transformation happens, sitting beneath the frameworks, libraries, and applications used across AI, deep learning, and HPC, as well as graphics, automotive, robotics, and other CUDA-powered products. That gives this team unusual leverage: reduce overhead in a fundamental launch, synchronization, memory, or data-movement path-or create a new driver or runtime capability-and the improvement can flow through many downstream systems and be repeated across vast numbers of products. One well-designed systems feature can help customers obtain more useful work from GPUs already deployed while informing how future CUDA capabilities and GPU architectures are designed.

We are looking for systems software engineers who want to work at this leverage point. You will design and ship production C/C++ features and optimizations in the CUDA driver and runtime, trace important workloads across application, operating-system, CPU, interconnect, and GPU boundaries, bring up new platforms, and turn evidence into future software and hardware direction. Your work will not end at a benchmark: it can make AI tools more responsive and efficient, help scientists reach answers sooner, and enable intelligent machines and interactive products to operate within demanding real-time constraints. Over time, you can grow from owning critical features and performance paths to setting subsystem direction and leading hardware/software co-design across generations-helping build the computing foundation for the next decade of AI and accelerated computing.

What you'll be doing:

  • Design, implement, validate, and ship performance-centric features and programming-model capabilities in the CUDA driver and runtime, writing maintainable, well-tested production C/C++.

  • Optimize critical execution paths-including kernel launch, synchronization, memory management & movement, CPU-GPU coordination, and system interconnect use-for latency, throughput, bandwidth, efficiency, and scalability.

  • Own complex performance problems end-to-end - understand important workloads, form hypotheses, create focused measurements and models, isolate root causes across software and hardware boundaries, implement production solutions, and validate application-level impact.

  • Establish performance expectations for current and future platforms, characterize new silicon, close software and hardware gaps, and drive performance readiness through product release.

  • Translate workload and platform evidence into CUDA API and programming-model improvements, systems-software direction, and measurement-backed recommendations for future hardware architecture and implementation.

  • Partner with application, library, framework, operating-system, driver, runtime, firmware, GPU architecture, silicon, product, and customer-facing teams; communicate findings clearly and raise engineering quality through design and code reviews.

  • Lead complex feature development and cross-layer investigations across teams, define performance requirements and technical direction for major subsystems, mentor engineers, and shape hardware/software decisions for future product generations.

What we need to see:

  • A BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or a related field-or equivalent practical experience - with at least 2 years of relevant systems-software development experience.

  • Strong production C/C++ systems-programming experience, including delivery of substantial features, optimizations, or production fixes in a complex codebase.

  • Strong operating systems and concurrency foundations, including threads, synchronization, processes, virtual memory, and user/kernel interactions.

  • Strong computer-architecture foundations, including processors, memory hierarchy, caching and coherence, data movement, and system interconnects.

  • Demonstrated success improving real software performance: measuring behavior, identifying the limiting mechanism, implementing an effective solution, and validating the result quantitatively.

  • Sound technical judgment, ownership of ambiguous problems, and clear communication across organizational and disciplinary boundaries.

  • Direct CUDA or GPU experience is valuable but is not required when accompanied by deep systems-software, operating-systems, computer-architecture, and performance-engineering foundations.

Ways to stand out from the crowd:

  • Experience developing GPU or accelerator drivers, runtimes, kernel software, firmware, compilers, or other performance-critical low-level systems.

  • Experience with pre-silicon analysis, platform bring-up, performance modeling, or hardware/software co-design.

  • Systems-level performance experience with AI/DL, HPC, graphics, automotive, robotics, or similarly demanding workloads.

  • Evidence of technical invention(s), such as software-performance patents, novel production designs, or measurement-backed recommendations that influenced a hardware revision or future architecture.

  • Python or another scripting language used for focused experimentation, data analysis, or visualization.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 5, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,174,835 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$107k – $174k per year • In office • Full-Time • Bachelor's Degree • San Francisco
C++
Python
AI/ML
Computer Vision
Edge AI
Embodied AI
Multimodal AI
Reinforcement Learning
Self-Supervised Learning
Vision-Language-Action
Robotics
Imitation Learning
Motion Planning
Reinforcement Learning
ROS
ROS2
Sensor Fusion
SLAM
Apply
$20k – $51k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Noida
C++
JavaScript
Node JS
Python
DevOps
CI/CD
Cybersecurity
CVSS
CWE
OWASP ASVS
Threat Modeling
Apply
$85k – $242k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Rehovot
Java
Python
AI/ML
Computer Vision
DevOps
Docker
Kubernetes
Apply
$27k – $70k per year (Estimated) • Remote/Hybrid • 11+ years exp • Noida
JavaScript
Node JS
Python
SQL
TypeScript
Java
Java
Hibernate
Spring MVC
Databases
PostgreSQL
Frontend
JQuery
React.js
DevOps
Amazon CloudWatch
Amazon S3
AWS
AWS Lambda
CI/CD
Docker
IAM
Jenkins
Kubernetes
Platform Engineering
Rest API
Apply
In office • 10+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
DevOps
Azure
Analytics
ETL/ELT
Power BI
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
Agentic Workflows
AI Agents
Function Calling
LangChain
LlamaIndex
LLM
OpenAI
Tool Use
DevOps
CI/CD
Git
Apply
In office • Internship • PhD • Shanghai
C++
Python
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenCL
DevOps
HPC
Apply
In office • Internship • Master's Degree • Shanghai
C#
C++
Python
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
AI Agents
Apply
In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$171k – $324k per year (Estimated) • Equity • In office • 15+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Cybersecurity
Zero Trust
Marketing
Instagram
LinkedIn
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
See all jobs
This is one of many
1,174,835 more open roles from verified company boards, updated every day.