691,154open jobs
40,551companies
97,801added this week
Browse all
Salary
$208k – $328k per year
Location
In office (Santa Clara, New York, Seattle)
Seniority
Principal · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is driving a vision for AI factories that convert tokens to intelligence at scale to power AI demands of tomorrow. Maintaining AI infrastructure at scale takes more than human involvement; it demands smart automation.

We're hiring a Technical Product Manager to drive AI factory resilience platform features and developer experience. You'll build underpinnings and glue for resilience of AI Factories to make it more adaptable across all architectures, workloads and generations of hardware. If you are excited about being a part of a 0 --> 1 effort to create and establish an open source project on AI Infrastructure resilience we want to hear from you!

What You'll Be Doing:

  • Define the resilience platform, own the product roadmap and delivery for specific platform features -such as common telemetry interfaces, health-check contracts, attribution hooks, and observability APIs that all products conform to, joint deployment experience etc,

  • Developer pain into features and integration partnerships.

  • Collaborate with the open-source developer community - prioritizing GitHub issues, gathering feedback, supporting contributors, and channeling community signal back into the roadmap.

  • Collaborate with engineering on feature design, prioritization, execution, and architecture tradeoffs.

  • Align cross-functionally with other Product, Engineering, Product Marketing, and Field teams on requirements, roadmaps, messaging, and engagements.

What We Need To See:

  • 12+ years in product management, solutions architecture, or software engineering on a technical product.

  • Bachelor's degree in Computer Science or an equivalent experience.

  • Technical depth in 2 or more of Data center operations, GPU infra, network and storage, container orchestration, (Kubernetes) , developer platforms and SDKs, agent frameworks.

  • Proven capability to connect with senior technical customers and translate requirements into product strategy.

  • Comfort operating in fast paced environments.

  • Strong written and verbal communication across Developers to Executive audiences.

Ways To Stand Out From The Crowd:

  • Strong experience building products for Data Center infra operations and observability

  • Practical experience delivering or contributing to an open-source product, including interacting with contributors on GitHub.

  • Experience crafting developer-facing APIs, SDKs, or CLIs at scale.

  • Background as an SRE or building SRE focused products

NVIDIA is widely considered one of the technology world’s most desirable employers. We have some of the world's most forward-thinking and hardworking people on our team. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 327,750 USD for Level 5, and 240,000 USD - 379,500 USD for Level 6.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 22, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
691,154 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$97k – $108k per year • Remote/Hybrid • Confidential • Full-Time • Newport • Cardiff
Java
Java
Liquibase
Databases
Google BigQuery
Google Cloud Spanner
BigQuery
DevOps
Terraform
GCP
Istio
Dynatrace
CI/CD
Jenkins
Kubernetes
Platform Engineering
Nexus Repository
GitHub
IAM
Cybersecurity
SonarQube
Management
Agile
Apply
$68k – $157k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Atlanta
Python
JavaScript
TypeScript
SQL
Python
pySpark
Databases
Snowflake
ElasticSearch
AI/ML
Hadoop
Spark
Amazon SageMaker
DevOps
AWS
Kubernetes
Apply
In office • Full-Time • Bachelor's Degree • Courbevoie
Python
Bash
AI/ML
CUDA Toolkit
CUDA
InfiniBand
NVLink
DevOps
Terraform
Ansible
Red Hat
Loki
Prometheus
SLURM
CI/CD
GitOps
Kubernetes
Ubuntu
Grafana
Configuration Management
KubeVirt
SLI/SLO/SLA
HPC
Linux
Apply
$28k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Pune
JavaScript
SQL
C#
Visual Basic
C#
ASP.NET Core
WPF
Databases
Db2
AI/ML
Copilot
Prompt Engineering
RAG
OpenAI
Mobile
MAUI
DevOps
Azure DevOps
Azure
CI/CD
GitHub
Apply
$28k – $64k per year (Estimated) • In office • Full-Time • 10+ years exp • Pune
JavaScript
Node JS
DevOps
Git
AWS
Kubernetes
Apply
OEM Sales Director 2 hours ago
$296k – $449k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
Apply
$132k – $207k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Python
AI/ML
InfiniBand
DevOps
HPC
Linux
Apply
$124k – $196k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin
Python
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Management
Outlook
Microsoft Office
Apply
$124k – $227k per year (Estimated) • Remote • Full-Time • 8+ years exp • Australia
Go
AI/ML
Edge AI
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • PhD • Santa Clara • Austin • Redmond
Python
Java
AI/ML
AI Agents
Agentic Workflows
Cybersecurity
PKI
Apply
In office • Full-Time • Bachelor's Degree • Santa Clara
Apply
OEM Sales Director 2 hours ago
$296k – $449k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Management
Outlook
Microsoft Office
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
Quantization
TensorRT
PyTorch
CUDA
Machine Learning
DevOps
Linux
Apply
See all jobs
This is one of many
691,154 more open roles from verified company boards, updated every day.