402,911open jobs
14,044companies
78,108added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA’s accelerated computing platform is foundational to modern HPC and AI. At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software.

We are hiring a Senior Software Engineer to advance the C++ foundation of CUDA Core Libraries. You will design and optimize high-performance algorithms and APIs for C++ developers. You will join the team building the foundational libraries, algorithms, and language/runtime infrastructure that make CUDA a speed-of-light experience for developers and AI coding agents alike.

What you’ll be doing:

  • Design and implement foundational CUDA C++ libraries, parallel algorithms, utilities, and runtime abstractions.

  • Compose and optimize GPU algorithms from high-level generic interfaces through low-level implementation.

  • Design stable interoperability boundaries that allow core C/C++ functionality to be consumed efficiently from Python and Rust.

  • Balance performance, compile time, portability, compatibility, usability, and long-term API evolution.

  • Own features throughout their lifecycle: design, implementation, testing, profiling, benchmarking, documentation, release, and maintenance.

  • Improve developer productivity through diagnostics, examples, build integration, tests, benchmarks, and continuous integration.

  • Collaborate with Python, Rust, compiler, and runtime engineers during architecture, design, and code reviews.

  • Engage with users on performance investigations, API feedback, and correctness issues.

What we need to see:

  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience and 8+ years of relevant software-development experience.

  • Strong production programming skills in C and C++, with deep knowledge of modern C++.

  • Experience with generic programming, templates, type systems, and standard-library design principles.

  • Proven experience developing systems-level software with demanding performance, concurrency, and compatibility requirements.

  • Practical experience with CUDA or another parallel or heterogeneous programming environment.

  • Experience developing production software or foundational libraries, including testing, profiling, benchmarking, and code review.

  • Understanding of API and ABI compatibility and the challenges of exposing C/C++ functionality to other languages.

  • Ability to work independently, define project scope, and drive complex work to completion.

  • Clear written communication skills for architecture documents, API specifications, and developer documentation.

  • Comfort working in large C/C++ codebases with build systems, toolchains, and continuous-integration infrastructure.

Ways to stand out from the crowd:

  • Strong understanding of CPU/GPU architecture and performance optimization, with hands-on experience in GPU-accelerated stacks (CUDA C++/Python, PyTorch, JAX, Numba, CuPy, or similar).

  • Proficiency with modern C++ and GPU libraries such as Thrust, CUB, and libcudacxx.

  • Experience with compiler infrastructure and tooling, including LLVM, Clang, or MLIR.

  • Knowledge of binary interfaces, linking, versioning, cross-platform distribution, and interoperability across Python, Rust, and C/C++ stacks.

  • Demonstrated interest in developer tools, library design, and improving developer productivity.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 4, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
402,911 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$26k – $51k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Pune
Python
Scala
SQL
Python
pySpark
Databases
Apache Kafka
Kafka
AI/ML
Spark
DevOps
CI/CD
Docker
Git
Kubernetes
Platform Engineering
Apply
$21k – $52k per year (Estimated) • In office • 4+ years exp • Bengaluru
Apex
JavaScript
Apex
Lightning Web Components
AI/ML
Agentforce
DevOps
CI/CD
Git
Marketing
Salesforce
Apply
$21k – $50k per year (Estimated) • In office • Bachelor's Degree • Moscow
C++
Python
SQL
Databases
PostgreSQL
Trino
AI/ML
Airflow
Spark
Apply
Remote/Hybrid • Full-Time • Helsinki
C++
JavaScript
Python
SQL
TypeScript
DevOps
AWS
Azure
CloudFormation
GCP
Terraform
Apply
$30k – $53k per year (Estimated) • Remote • Saint Petersburg
C#
C#
.NET
Mobile
Firebase
DevOps
CI/CD
TeamCity
Apply
$30k – $79k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Pune
Bash
C++
Python
AI/ML
Claude
Claude Code
CUDA
CUDA Toolkit
Cursor
OpenAI Codex
Triton
DevOps
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Apply
$168k – $270k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Bash
C++
Python
AI/ML
CUDA
CUDA Toolkit
DevOps
CI/CD
GitLab
Jenkins
Apply
$12k – $47k per year (Estimated) • In office • Full-Time • 2+ years exp • Pune
C++
Python
AI/ML
AI Agents
Reinforcement Learning
DevOps
CI/CD
Robotics
Isaac Lab
Isaac Sim
Perception
Reinforcement Learning
ROS
Sim-to-Real
Apply
$121k – $285k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Australia
Bash
Python
AI/ML
InfiniBand
NCCL
NVLink
DevOps
Amazon S3
Ansible
Apply
$121k – $291k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Australia
Python
AI/ML
InfiniBand
DevOps
Ansible
CI/CD
HPC
SLURM
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
$144k – $230k per year • In office • Full-Time • PhD • Santa Clara
Apex
Apex
MuleSoft
Apply
See all jobs
This is one of many
402,911 more open roles from verified company boards, updated every day.