986,892open jobs
58,977companies
161,835added this week
Browse all
Salary
$116k – $190k per year
Location
In office (Santa Clara, Austin, United States, Redmond)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is at the forefront of the generative AI revolution, building the software and systems that power the world’s most advanced large language model workloads. We are looking for a Software Engineer focused on bring-up, triage, benchmarking, analysis, and optimization of distributed training and inference workloads across NVIDIA GPU platforms at the largest scales we run.

In this role you will help bring up, benchmark, and debug distributed LLM workloads on multi-GPU and multi-node deployments, and own the design and implementation of the benchmarking tooling, automation, and debugging workflows that support them. This is a hands-on role for an engineer who enjoys deep technical problems across deep learning systems, GPU performance, distributed computing, and large-scale operations.

What you’ll be doing:

  • Bring up, validate, and debug large-scale AI clusters, infrastructure, and end-to-end workloads.

  • Bring up, tune, and benchmark AI pre-training, post-training, and inference workloads using PyTorch, NeMo / Megatron, TensorRT-LLM, and adjacent NVIDIA AI software stacks.

  • Perform root-cause analysis of failures in large distributed environments

  • Contribute to the resilience and failure-attribution tooling that detects, triages, and attributes node, fabric, and workload failures across the cluster.

  • Build and maintain repeatable benchmark suites, automation, acceptance criteria, and qualification workflows on new platforms.

  • Tune runtime settings, communication parameters, and deployment configurations in close partnership with framework, systems, and platform teams.

  • Deliver actionable, data-driven recommendations based on profiling, benchmark results, and cluster characterization.

What we need to see:

  • Bachelor’s or Master’s in Computer Science or a related technical field (or equivalent experience).

  • 3+ years of experience developing software for AI, HPC, or systems-level applications.

  • Hands-on experience with multi-GPU or multi-node workloads and CUDA-aware distributed execution.

  • Backgroun with debugging and scaling distributed systems.

  • Experience debugging and triaging AI applications across the full stack, from the application level toward the hardware.

  • Experience operating workloads in scheduled, containerized cluster environments.

  • Excellent analytical, debugging, and communication skills, and a collaborative approach across teams.

  • Strong Python and C/C++ programming skills.

Ways to stand out from the crowd:

  • Hands-on experience with NCCL and CUDA-aware distributed execution.

  • Deep familiarity with the RDMA software stack (NCCL, IB verbs, UCX, libfabric) and with InfiniBand / RoCE congestion debugging.

  • Experience building acceptance tests, benchmark harnesses, regression gates, or cluster qualification tooling for AI platforms, including MLPerf.

  • Experience diagnosing performance jitter

  • Experience building resilience, fault-detection, or failure-attribution systems for datacenter-scale infrastructure.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you’re creative, autonomous, and love a challenge, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 189,750 USD for Level 2, and 140,000 USD - 224,250 USD for Level 3.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 3, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
986,892 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Santa Clara
Software Engineer I 8 hours ago
$77k – $140k per year • Remote (United States) • 1+ year exp • Bachelor's Degree
AI/ML
Copilot
Cursor
Management
Agile
Apply
$140k – $200k per year • Hybrid • TS/SCI • Full-Time • Arlington
AI/ML
LangGraph
LangChain
Prompt Engineering
AI Agents
Pydantic AI
LLM
LLM Guardrails
DevOps
CI/CD
Management
Agile
Apply
≈ $106k – $214k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Jersey City
Python
Databases
PostgreSQL
OpenSearch
AI/ML
Fine-tuning
Prompt Engineering
Function Calling
AI Agents
AWS Bedrock
LLM
RAG
Agentic Workflows
Tool Use
Machine Learning
DevOps
AWS
Amazon ECS
Management
Agile
Apply
$150k – $170k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Arlington
Python
Java
SQL
C#
C#
.NET
Databases
MS SQL
DevOps
CloudFormation
Azure
CI/CD
Jenkins
Git
AWS
Bitbucket
Cybersecurity
SonarQube
Analytics
Power BI
ETL/ELT
Talend
Management
Confluence
Jira
Microsoft Teams
Agile
Scrum
Apply
$111k – $172k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bellevue
AI/ML
Copilot
ChatGPT
Apply
$16k per year (net) • In office • Moscow
Python
JavaScript
PHP
SQL
C#
C++
DevOps
Rest API
Debian
CI/CD
GitLab
Linux
Astra Linux
TCP/IP
Cybersecurity
OWASP Top 10
Apply
In office • 2+ years exp • Bachelor's Degree • Minsk
Python
SQL
Databases
PostgreSQL
ClickHouse
Apache Kafka
AI/ML
Polars
MLFlow
XGBoost
LightGBM
NumPy
PyTorch
Anomaly Detection
Time Series Forecasting
DevOps
CI/CD
Docker
Apply
≈ $76k – $136k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Manching
Python
Go
C++
C++
Qt
DevOps
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Management
Confluence
Jira
Agile
Apply
≈ $49k – $112k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Istanbul
Python
SQL
Python
pySpark
AI/ML
Spark
dbt
Machine Learning
Analytics
Tableau
Power BI
Superset
Cognos
Apply
In office • Full-Time • Master's Degree • Warsaw
Python
SAS
AI/ML
Machine Learning
Management
Agile
Scrum
Kanban
Apply
$152k – $242k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Santa Clara • Westford • Saint Louis • Boulder • New York
Python
Bash
DevOps
NixOS
Debian
CI/CD
Docker
Kubernetes
Ubuntu
HPC
Linux
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara • Durham
Apply
$152k – $242k per year • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree • United States
Go
Rust
C++
Databases
ClickHouse
AI/ML
CUDA Toolkit
CUDA
Frontend
GraphQL
DevOps
Datadog
Kubernetes
Grafana
HPC
Apply
$184k – $288k per year • Remote (United States) • Full-Time • 6+ years exp • Bachelor's Degree • United States
DevOps
Helm
ArgoCD
Kubernetes
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • Bachelor's Degree • Durham
Python
DevOps
SLURM
Kubernetes
Incident Management
Apply
$180k – $270k per year • In office • Santa Clara
Python
C++
DevOps
Linux
Apply
$207k – $280k per year • Equity • In office • Full-Time • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
Apply
$45k – $100k per year • In office • 1+ year exp • Santa Clara
Apply
$71k – $79k per year • In office • Santa Clara
Apply
$71k – $79k per year • In office • Santa Clara
Apply
See all jobs
This is one of many
986,892 more open roles from verified company boards, updated every day.