1,443,753open jobs
85,355companies
221,329added this week
Browse all
Salary
$320k – $489k per year
Location
In office (Santa Clara)
Seniority
Architect · 7+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 9, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking an engineering director to lead the software teams behind our bare-metal datacenters and cloud compute infrastructure. The organization supports hundreds of megawatts of datacenter capacity already online, with more capacity coming. You will build the software that makes this growing physical infrastructure available as reliable, secure, and efficient compute services for NVIDIA engineering.

Our team provides NVIDIA’s continuous integration (CI) infrastructure: the environment where hardware, firmware, drivers, networking, and system software are integrated and tested on their path to production. You will enable engineering teams to bring up preproduction systems, reproduce failures, validate changes, and move platforms toward production readiness. The fleet spans multiple hardware generations and maturity levels: x86 and Arm servers, GPUs, DPUs, high-speed networking, storage, and rack-scale, liquid-cooled systems. Platforms such as Grace Blackwell and Vera Rubin illustrate the breadth of compute and interconnect technology involved. This role combines hands-on systems judgment with leadership of the teams making that diversity manageable at scale.

What you’ll be doing:

  • Lead software engineering teams responsible for cloud compute, bare-metal fleet management, and CI infrastructure, owning architecture, implementation, deployment, and operations.
  • Set the technical direction for compute control planes, resource provisioning, topology-aware placement, reservations, and capacity management across physical servers, virtual machines, and containers.
  • Automate the bare-metal lifecycle: hardware discovery and inventory, server bring-up, imaging, firmware and driver configuration, out-of-band management, health checks, reprovisioning, and recovery.
  • Build CI services that allocate the right hardware, provision repeatable test environments, record hardware and software configurations, collect diagnostics, and restore systems to a known state. Shorten the path from a hardware or software change to actionable validation results.
  • Keep engineering services dependable while supporting evolving preproduction hardware and software. Define service-level objectives, isolate failures, improve observability, and turn incidents and recurring test-infrastructure failures into engineering fixes.
  • Measure and improve hardware onboarding time, provisioning speed, CI queue time, productive fleet utilization, recovery time, and cost efficiency as capacity expands.
  • Set engineering standards for design and code reviews, automated testing, secure development, release quality, and safe changes to production systems.
  • Hire, coach, and retain engineers and engineering managers. Establish clear ownership, develop technical leaders, and build teams that deliver consistently over multiple release cycles.
  • Partner with hardware, firmware, drivers, networking, storage, security, and validation teams to onboard platforms and diagnose failures across system boundaries. Translate engineering users’ needs into clear priorities and technical decisions.
  • Work with datacenter operations and facilities teams to bring additional capacity online. Connect rack power, cooling, physical topology, and hardware-health telemetry to provisioning, placement, serviceability, and operational readiness.

What we need to see:

  • 15+ overall years of experience in software engineering, distributed systems, or cloud infrastructure, including 7+ years leading engineering teams and experience managing engineering managers.
  • Direct engineering ownership of the underlying services of a public or private compute cloud, such as AWS EC2, Google Compute Engine, Azure Compute, OCI Compute, or a comparable infrastructure-as-a-service platform.
  • Strong technical depth in distributed systems, Linux, virtualization, containers, and the networking and storage services that support large compute fleets.
  • Experience building software and automation for bare-metal infrastructure, including server provisioning, hardware inventory, firmware or operating-system lifecycle management, and recovery across heterogeneous systems.
  • Experience running production infrastructure with demanding availability requirements, including failure isolation, incident response, observability, and reliable deployment and recovery mechanisms.
  • The ability to guide architecture, evaluate implementation choices, and debug complex interactions among hardware, firmware, drivers, operating systems, and distributed services with senior engineers.
  • A record of sustained ownership as platforms evolve, with measurable improvements in reliability, scalability, or engineering delivery. Strong communication, collaboration, and people-development skills.
  • A degree in computer science, computer engineering, or a related discipline, or equivalent experience.

Ways to stand out from the crowd:

  • Experience building or operating OpenStack infrastructure, particularly Nova, Neutron, Cinder, or Ironic, or contributing to related open-source projects.
  • Engineering leadership spanning compute control planes and site reliability, including multi-tenant isolation, scheduling, placement, and capacity allocation.
  • Experience with GPU clusters, rack-scale computing, NVLink, InfiniBand or high-speed Ethernet, and AI or high-performance computing workloads.
  • Experience supporting preproduction hardware, new product introduction, or hardware-in-the-loop CI, including repeatable validation environments and systematic regression isolation.
  • Familiarity with liquid-cooled, high-density datacenters and how power, thermal constraints, cooling, and component health affect fleet availability and scheduling and delivery of compute platforms across multiple regions, datacenters, or hybrid-cloud environments, including fleet upgrades and workload migration with minimal service disruption.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,443,753 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Leadership
Similar stack
Same company
Santa Clara
$228k – $285k per year • In office • 7+ years exp • Bachelor's Degree • Palo Alto
Python
C++
AI/ML
TensorRT
Apply
$265k – $331k per year • Equity • In office • 7+ years exp • Bachelor's Degree • Palo Alto
Python
AI/ML
Quantization
TensorRT
Machine Learning
Apply
Engineering Manager 1 hour ago
$142k – $214k per year • In office • 8+ years exp • Las Vegas
Apply
$142k – $214k per year • Equity • In office • 10+ years exp • Bachelor's Degree • New Haven
Management
Microsoft Office
Apply
≈ $146k – $275k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Morrisville
DevOps
GCP
Azure
AWS
Cybersecurity
ISO 27001
NIST CSF
PCI DSS
SOC 2
GDPR
HIPAA
Threat Modeling
Apply
$118k – $129k per year • Remote (United States, PST hours) • Full-Time • 5+ years exp
DevOps
Azure
CI/CD
Windows Server
Git
AWS
GitHub
VPN
VLAN
Cybersecurity
Microsoft Defender
Microsoft Entra ID
Apply
≈ $73k – $144k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Scottsdale
Python
Java
SQL
Java
Spring Framework
Databases
DynamoDB
AI/ML
Claude
ChatGPT
Prompt Engineering
DevOps
Rest API
Terraform
CloudFormation
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
AWS Lambda
Amazon EC2
Amazon S3
IAM
AWS Step Functions
API Gateway
Apply
≈ $72k – $167k per year (Estimated) • Hybrid • PhD • Montreal
Python
SQL
Python
FastAPI
AI/ML
LangChain
Scikit-learn
TensorFlow
PyTorch
LLM
RAG
LLMOps
DevOps
GCP
Azure DevOps
Azure
CI/CD
AWS
FinOps
Apply
Software Pentester 1 day ago
≈ $39k – $69k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • Poland
Python
JavaScript
PHP
C#
Perl
DevOps
Puppet
Kali Linux
Chef
CI/CD
AWS
Docker
Kubernetes
Linux
TCP/IP
Cybersecurity
Nmap
Nessus
Checkmarx
CWE
Fortify
LDAP
SIEM
OWASP
Robotics
Digital Twin
Management
Agile
Apply
$160k – $175k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Irving
SQL
Databases
Pinecone
ElasticSearch
AI/ML
Triton Inference Server
Machine Learning
DevOps
Terraform
GCP
Azure
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Management
Agile
Waterfall
Apply
$320k – $483k per year • In office • Full-Time • 18+ years exp • Bachelor's Degree • Santa Clara
Apply
$168k – $265k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Apply
$168k – $270k per year • Remote (United States) • Full-Time • 3+ years exp • Bachelor's Degree • United States
Go
Java
Rust
C++
Frontend
GraphQL
DevOps
Rest API
Terraform
Helm
OpenTofu
WebRTC
containerd
CRI-O
GitOps
ArgoCD
AWS
Docker
Kubernetes
Amazon EKS
AWS Fargate
IAM
Amazon ECS
Apply
$208k – $334k per year • In office • Full-Time • 10+ years exp • PhD • Santa Clara
Python
Go
JavaScript
TypeScript
AI/ML
AI Agents
Agentic Workflows
DevOps
GCP
OpenTelemetry
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Incident Management
Linux
Apply
$224k – $357k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
RAG
Apply
$48k per year • In office • Full-Time • High School Diploma • Santa Clara
Management
Agile
Apply
$102k – $210k per year • Equity • In office • 10+ years exp • Bachelor's Degree • Santa Clara
DevOps
GCP
Azure
AWS
Analytics
Microsoft Excel
Apply
$140k – $200k per year • In office • Full-Time • 8+ years exp • Santa Clara
MATLAB
DevOps
Red Hat
VMWare
SLURM
Azure
CI/CD
Jenkins
AWS
Ubuntu
KVM
Xen
HPC
Linux
Windows
Cybersecurity
Active Directory
LDAP
Apply
$213k – $288k per year • Equity • In office • Full-Time • 7+ years exp • Santa Clara
AI/ML
vLLM
Reinforcement Learning
Computer Vision
AI Agents
TensorFlow
PyTorch
Amazon SageMaker
SFT
Post-training
Megatron-LM
FSDP
TPU
Machine Learning
Apply
$150k – $262k per year • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
SQL
Databases
PostgreSQL
Redis
NATS
RabbitMQ
Apache Kafka
Amazon Aurora
AI/ML
AI Agents
DevOps
Terraform
GCP
Crossplane
etcd
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Amazon EKS
Google GKE
Azure AKS
IAM
Apply
See all jobs
This is one of many
1,443,753 more open roles from verified company boards, updated every day.