1,454,406open jobs
86,438companies
225,454added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 9, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA DGX Cloud builds and operates large-scale GPU infrastructure for AI workloads. We are looking for Software Engineers with SRE or Production Engineering experience who have worked hands-on with bare-metal NVIDIA systems. This team builds the software and operational tooling that moves GPU capacity from installed hardware to production service supporting an IaaS production environment of BMaaS, VMaaS.

What makes this opportunity outstanding is the chance to work with innovative technology to develop the future of AI computing. Join us to be part of a world-class team and make an impact on the next era of computing! At NVIDIA, you’ll help make next-generation AI infrastructure production-ready at scale!

What you’ll be doing:

  • Build automation for bare-metal provisioning, hardware validation, firmware and software upgrades, repair, and cluster lifecycle management.
  • Build tools using BMC and Redfish interfaces to assess hardware health, regulate server state, and facilitate recovery workflows.
  • Manage and enhance NVIDIA NVL72 systems and BlueField-3 or later DPUs within cloud partner and on-premises environments.
  • Diagnose failures across servers, DPUs, GPU systems, CPU systems, networking, Linux, and Kubernetes; turn recurring issues into automated detection and repair.
  • Define validation and handoff criteria so new capacity enters production safely and consistently.
  • Take part in on-call duties, incident response, root-cause analysis, and ensure permanent resolutions are implemented.
  • Collaborate with hardware, networking, platform, data center operations, and partner teams to resolve issues across ownership boundaries.

What we need to see:

  • 8+ years building software for or operating production infrastructure, including substantial hands-on bare-metal experience.
  • Strong Go or Python skills, with a record of delivering production automation and services.
  • Direct experience working with BMC and Redfish for server provisioning, health inspection, power control, or fault diagnosis.
  • Practical experience working directly with NVIDIA GPU hardware, including NVL72 systems, and BlueField-3 or later DPUs.
  • Experience with Linux, firmware and driver management, network boot, and the server lifecycle from initial provisioning through repair.
  • Experience managing production reliability via on-call duties, incident handling, observability, and durable solutions.
  • Ability to debug failures across hardware, host operating systems, networking, and distributed services.
  • Clear communication and demonstrated ownership of problems that span multiple teams.
  • BS/MS in Computer Science or equivalent experience in a related field.

Ways to stand out from the crowd:

  • Experience operating BlueField DPUs in DPU mode, including host-to-DPU connectivity and lifecycle debugging, or equivalent experience.
  • Background with NVLink, InfiniBand, Spectrum-X, or GPU cluster performance validation.
  • Experience building safe, repeatable workflows for rack-scale bringup, firmware upgrades, hardware replacement, and customer handoff.
  • Background with Kubernetes, GitOps, Argo CD, SLOs, and fleet-wide automation.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,454,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Santa Clara
≈ $115k – $223k per year (Estimated) • In office • Full-Time • 4+ years exp • Charlotte
PowerShell
DevOps
Azure
Windows
Wi-Fi
Cybersecurity
Microsoft Entra ID
Active Directory
Management
ServiceNow
Microsoft Office
Apply
$156k – $261k per year • In office • 10+ years exp • Bachelor's Degree • Santa Rosa
Python
Java
C#
C++
Perl
MATLAB
Apply
≈ $94k – $182k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Chandler
Python
Bash
Perl
DevOps
Puppet
Ansible
Red Hat
VMWare
Chef
Prometheus
Git
Docker
Kubernetes
Grafana
Configuration Management
Nagios
Proxmox VE
Incident Management
SLI/SLO/SLA
HPC
Linux
Unix
Apply
$120k – $199k per year • Equity • In office • 4+ years exp • Bachelor's Degree • San Diego
Python
C++
DevOps
Linux
Apply
$125k – $167k per year • Equity • In office • 6+ years exp • Bachelor's Degree • Austin
Python
C++
DevOps
Linux
Apply
$140k – $200k per year • In office • Full-Time • 8+ years exp • Santa Clara
MATLAB
DevOps
Red Hat
VMWare
SLURM
Azure
CI/CD
Jenkins
AWS
Ubuntu
KVM
Xen
HPC
Linux
Windows
Cybersecurity
Active Directory
LDAP
Apply
$48k per year • In office • Full-Time • High School Diploma • Santa Clara
Management
Agile
Apply
$102k – $210k per year • Equity • In office • 10+ years exp • Bachelor's Degree • Santa Clara
DevOps
GCP
Azure
AWS
Analytics
Microsoft Excel
Apply
$213k – $288k per year • Equity • In office • Full-Time • 7+ years exp • Santa Clara
AI/ML
vLLM
Reinforcement Learning
Computer Vision
AI Agents
TensorFlow
PyTorch
Amazon SageMaker
SFT
Post-training
Megatron-LM
FSDP
TPU
Machine Learning
Apply
$150k – $262k per year • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
SQL
Databases
PostgreSQL
Redis
NATS
RabbitMQ
Apache Kafka
Amazon Aurora
AI/ML
AI Agents
DevOps
Terraform
GCP
Crossplane
etcd
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Amazon EKS
Google GKE
Azure AKS
IAM
Apply
See all jobs
This is one of many
1,454,406 more open roles from verified company boards, updated every day.