982,073open jobs
58,789companies
159,182added this week
Browse all
Salary
$272k – $431k per year
Location
In office (Santa Clara, Seattle)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Sep 29, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

The NVIDIA Kubernetes Engine (NKE) team is looking for a technical leader to lead the Runtime Engineering team responsible for the full configuration lifecycle of NKE tenant workload clusters. This team is responsible for software components that keep GPU workloads reliable and secure at scale. Their scope includes cluster bootstrapping, node configuration, and the container execution environment, including NVIDIA's AI Container Runtime (AICR). You will work across networking, storage, GPU resource management, and cluster security to deliver a production-grade, multi-tenant Kubernetes platform. Your team's decisions directly shape the runtime foundation that internal and external customers depend on.

What You'll Be Doing:

  • Be responsible for the build, implementation, and operational reliability of cluster configurations for NKE tenant workloads across all supported topologies
  • Manage a team of engineers coordinating the entire container runtime stack: AICR, GPU management operator, DCGM, and related node-level components
  • Drive architecture decisions for cluster networking (CNI), storage (CSI), cluster HA , and GPU resource partitioning (MIG, MPS, time-slicing)
  • Define and implement cluster hardening standards, RBAC models, pod security policies, and multi-tenancy isolation boundaries
  • Partner with NKE platform, infrastructure, and cybersecurity teams to integrate new capabilities and resolve cross-cutting runtime concerns
  • Build and maintain tooling for AICR lifecycle management - provisioning, upgrades, configuration drift detection, and remediation
  • Represent the runtime team in architecture reviews, roadmap planning, and customer communications with NVIDIA leadership
  • Contribute to open source communities anywhere NKE has upstream dependencies or influence

What We Need to See:

  • BS/MS degree in Computer Science or related field (or equivalent experience)
  • 12+ overall years of relevant experience designing and delivering large-scale distributed software systems, including 5+ years of people-management experience leading, developing, and scaling high-performing software engineering teams responsible for complex, production-critical software.
  • Experience leading a group of engineers with varying specializations and seniority levels - bridging runtime, networking, and security fields is a core part of this role
  • Kubernetes internals knowledge - not just usage; you understand how the scheduler, kubelet, API server, and admission controllers interact
  • Cluster lifecycle management experience - Cluster API, kubeadm, or equivalent; experience leading fleet-scale cluster provisioning and upgrades
  • Security and compliance posture - CIS Kubernetes Benchmark, pod security admission, image signing, supply chain integrity
  • Proven ability to design and implement maintainable APIs for consumers
  • Familiarity with Identity and Access Management approaches
  • Excel in managing up, down, and across organizations
  • Demonstrated ability to reach cross-organization consensus without all the details

Ways to Stand Out from the crowd:

  • Prior experience with NVIDIA GPU Operator, DCGM Exporter, or NVLink-aware scheduling
  • Experience running Kubernetes at hyperscale with GPU node pools
  • Track record of upstream open source contributions in the Kubernetes or any open source runtime ecosystem
  • Experienced, persuasive, and effective interpersonal skills - written, verbal, and in front of engineering leadership
  • Demonstrated skills in coaching, analysis, problem solving, and short/long-term technical planning

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction - from artificial intelligence to autonomous vehicles. NVIDIA is widely considered one of the technology world's most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. If you're passionate about building the infrastructure that runs AI at scale, we want to hear from you.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 3, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
982,073 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Santa Clara
≈ $129k – $251k per year (Estimated) • In office • Top Secret • 2+ years exp • Bachelor's Degree • United States
Python
PowerShell
C
C
MPI
Databases
Apache Kafka
Amazon Redshift
AI/ML
Spark
CUDA Toolkit
CUDA
cuDNN
DevOps
SLURM
CI/CD
GitOps
AWS
HPC
Management
Agile
Apply
≈ $132k – $257k per year (Estimated) • In office • Top Secret • 8+ years exp • Bachelor's Degree • Chantilly
Python
C#
C#
.NET
Databases
DynamoDB
ElasticSearch
OpenSearch
Amazon Aurora
DevOps
Splunk
Terraform
Ansible
GCP
Red Hat
OpenShift
CloudFormation
GitLab CI
Azure
CI/CD
AWS
Docker
Kubernetes
Ubuntu
Bicep
AWS Lambda
API Gateway
HPC
Linux
Windows
Cybersecurity
Keycloak
Microsoft Entra ID
Active Directory
LDAP
Apply
Platform Architect 1 day ago
≈ $92k – $178k per year (Estimated) • Remote (United States) • Full-Time • 7+ years exp • Bachelor's Degree • United States
Python
TypeScript
Databases
DynamoDB
OpenSearch
AI/ML
AWS Bedrock
Amazon SageMaker
DevOps
AWS CDK
Pulumi
Git
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon EC2
FinOps
Amazon S3
IAM
Amazon ECS
Cybersecurity
Checkov
SOC 2
HIPAA
Apply
$96k – $120k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Oakland
DevOps
Azure
Windows Server
Windows
DNS
DHCP
VPN
VLAN
Cybersecurity
FortiGate
Microsoft Entra ID
Active Directory
Management
ITSM
Apply
≈ $112k – $223k per year (Estimated) • Hybrid • Secret • Full-Time • 10+ years exp • Bachelor's Degree • Washington
DevOps
Azure
Git
Cybersecurity
Microsoft Entra ID
SIEM
Management
Agile
Scrum
Apply
≈ $14k – $27k per year (Estimated) • In office • 1+ year exp • Krasnodar
Python
SQL
Databases
MySQL
PostgreSQL
Redis
RabbitMQ
DevOps
CI/CD
Docker
Kubernetes
Linux
Analytics
Power BI
Data Vault
Apply
≈ $46k – $84k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Rennes
Python
PowerShell
Bash
Databases
PostgreSQL
MinIO
DevOps
Terraform
Ansible
Helm
CI/CD
Kubernetes
Amazon S3
Linux
Apply
Software Engineer 12 hours ago
≈ $13k – $34k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Noida
Python
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Databases
MongoDB
Redis
GraphDB
DynamoDB
Apache Kafka
Frontend
Angular
React.js
Vite
DevOps
Rest API
GCP
Datadog
CI/CD
AWS
Docker
Kubernetes
Apply
≈ $45k – $82k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Gémenos
Databases
MariaDB
DevOps
Terraform
Ansible
Datadog
GitLab CI
CI/CD
Windows Server
Git
Docker
Kubernetes
GitLab
Windows
Apply
Senior Platform Lead 12 hours ago
$84k – $178k per year • In office • Full-Time • Bengaluru
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara • Austin • Redmond
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
JAX
TensorFlow
PyTorch
Ray
Post-training
Pre-training
NCCL
DevOps
Loki
Prometheus
Grafana
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
InfiniBand
NVLink
DevOps
GitOps
ArgoCD
Kubernetes
Linux
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Python
Perl
AI/ML
Cursor
DevOps
Linux
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara • Westford
Python
Rust
C++
AI/ML
AI Agents
InfiniBand
NVLink
DevOps
RTOS
Linux
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara • Hillsboro • Redmond
Python
C++
DevOps
Windows
Apply
$169k – $201k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
Python
PHP
SQL
Ruby
PowerShell
Bash
Perl
Databases
MySQL
Oracle
DevOps
Linux
Windows
Unix
Apply
GCP Data Engineer 1 day ago
$60k – $149k per year • In office • 5+ years exp • Santa Clara
DevOps
GCP
Apply
$175k – $265k per year • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Python
Bash
AI/ML
LLM
Anomaly Detection
InfiniBand
NVLink
DevOps
Splunk
Terraform
Ansible
GCP
Datadog
Prometheus
SLURM
Azure
CI/CD
AWS
Kubernetes
Grafana
FinOps
AIOps
HPC
Linux
Apply
$207k – $280k per year • Equity • In office • Full-Time • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
Apply
$45k – $100k per year • In office • 1+ year exp • Santa Clara
Apply
See all jobs
This is one of many
982,073 more open roles from verified company boards, updated every day.