931,896open jobs
56,942companies
155,467added this week
Browse all
Salary
$208k – $328k per year
Location
In office (Santa Clara)
Seniority
Senior · 12+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 28, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

At NVIDIA, we are building the foundation & blueprints for how AI workloads are served at scale for NVIDIA employees and DSX Partners. As a Product Manager for Inference Platform, you will help define and drive the products and platform capabilities that enable large-scale models serving across a broad portfolio of models, both open source and proprietary. You will work at the intersection of AI research, infrastructure engineering, and real user needs, shaping how inference is delivered reliably, efficiently, and at scale.

This is high-impact role. You will be expected to bring structure to undefined problem spaces, make progress without complete information, and build conviction through deep engagement with users, engineers, and the broader ecosystem of AI Cloud partners. The right candidate is equally comfortable discussing model serving architecture, token economics and writing a clear product brief.

What you will be doing:

  • Define product vision and strategy for inference platform capabilities - including APIs, capacity management, cost management, performance and optimization, and model serving infrastructure.
  • Translate user needs and infrastructure constraints into clear requirements and prioritized roadmaps.
  • Partner closely with engineering, research, and user groups teams to drive execution from concept through launch.
  • Be responsible for end-to-end product lifecycle for inference-related products and platform investments.
  • Develop deep understanding of the inference ecosystem - model formats, serving frameworks, API formats, and the tradeoffs that matter at scale.
  • Drive clarity in ambiguous situations by framing the problem, identifying what is known and unknown, and proposing a path forward.
  • Represent the voice of the user and ensure product decisions are grounded in real needs, not assumptions.
  • Track and synthesize developments across the inference landscape: open source model releases, serving frameworks, competitive dynamics, and emerging use cases.

What we need to see :

  • 12+ years of experience with a track record of delivering complex technical products.
  • Bachelors degree or higher, or equivalent experience
  • Strong written and verbal communication, you are able to write clear, concise product documents, specs, and strategies.
  • Ability to operate in ambiguous, fast paced environments and make progress without a full playbook.
  • Analytical professional who can break down complex problems, identify the right questions, and drive toward decisions.
  • Strong user empathy - able to synthesize qualitative and quantitative signals into a coherent picture of what users need and why.
  • Deep familiarity with AI/ML systems and inference serving frameworks (such as TensorRT-LLM, vLLM, or Triton Inference Server), tradeoffs involved, and what matters to model publishers, application developers and cloud operators.
  • Experience with inference APIs- design, versioning, performance, hardware efficiency, and developer experience.
  • Familiarity with open source as well as commercial model ecosystems and the different considerations each brings to a serving platform.

Ways to stand out from the crowd:

  • Experience with sophisticated inference techniques: disaggregated prefill/decode, KV cache management, speculative decoding, or continuous batching.
  • Background in large scale systems and developer platforms, cloud infrastructure, or MLOps tooling.
  • Exposure to capacity planning, quota management, or resource scheduling in large-scale compute environments.

You thrive in the early stages of building - where the problem is not fully defined, the team is still forming, and the decisions you make will shape direction for years. You don't wait for perfect information. You ask good questions, build conviction incrementally, and create momentum. You can balance user's reality and engineering constraints simultaneously. You bring clarity and cut through noise. You have a genuine curiosity about how inference works. You care about users’ needs beyond trivia, as it improves your product.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 208,000 USD - 327,750 USD for Level 5, and 240,000 USD - 379,500 USD for Level 6.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 2, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
931,896 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Product
Similar stack
Same company
Santa Clara
$208k – $340k per year • Hybrid • Full-Time • 8+ years exp • San Francisco • New York
Databases
Databricks
AI/ML
Embeddings
AI Agents
LLM
Apply
≈ $102k – $198k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Durham
AI/ML
Function Calling
AI Agents
LLM
RAG
Human-in-the-Loop
Multi-Agent Systems
Tool Use
Apply
Product Manager 2 months ago
≈ $99k – $182k per year (Estimated) • Remote (United States) • Full-Time • 3+ years exp • United States
Apply
≈ $142k – $259k per year (Estimated) • In office • 3+ years exp • Jersey City
Management
Agile
Apply
$245k – $280k per year • Remote (United States) • Full-Time • 5+ years exp
AI/ML
OpenRouter
LLMOps
LLM Guardrails
Cybersecurity
DLP
Apply
$94k – $131k per year (gross) • In office • Full-Time • 5+ years exp • Singapore
Python
Go
C++
AI/ML
vLLM
CUDA Toolkit
Triton Inference Server
Embeddings
Quantization
Multimodal AI
Function Calling
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
RAG
Reranking
NVIDIA NIM
CUDA
Hugging Face
NCCL
NVLink
cuDNN
Speculative Decoding
KV Cache
Agentic Workflows
Tool Use
DevOps
CI/CD
GitOps
Kubernetes
Apply
$270k – $310k per year • Hybrid • Full-Time • 8+ years exp • San Mateo • New York
AI/ML
vLLM
Fine-tuning
Multimodal AI
SGLang
TensorRT
TensorRT-LLM
PyTorch
LLM
Fireworks AI
DPO
SFT
DevOps
GCP
Azure
AWS
Apply
$224k – $357k per year • Remote (United States) • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Fine-tuning
Quantization
Knowledge Distillation
Diffusion Models
SGLang
Diffusers
TensorRT
TensorRT-LLM
PyTorch
CUDA
Triton
Edge AI
World Models
Model Distillation
Machine Learning
Game Dev
Unity
Unreal Engine
DLSS SDK
Apply
≈ $31k – $69k per year (Estimated) • Remote (EAEU) • 8+ years exp • Bachelor's Degree • Almaty
Python
AI/ML
Triton Inference Server
Fine-tuning
TensorRT
TensorRT-LLM
TensorFlow
PyTorch
LLM
RAG
RAPIDS
NVIDIA NIM
Triton
NVIDIA NeMo
LLM Guardrails
DevOps
Terraform
Ansible
GCP
Helm
VMWare
Azure
AWS
Kubernetes
Linux
Apply
≈ $103k – $231k per year (Estimated) • Remote (Georgia) • 8+ years exp • Bachelor's Degree • Tbilisi
Python
AI/ML
Triton Inference Server
Fine-tuning
TensorRT
TensorRT-LLM
TensorFlow
PyTorch
LLM
RAG
RAPIDS
NVIDIA NIM
Triton
NVIDIA NeMo
LLM Guardrails
DevOps
Terraform
Ansible
GCP
Helm
VMWare
Azure
AWS
Kubernetes
Linux
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • PhD • Santa Clara
AI/ML
Copilot
Cursor
Claude Code
AI Agents
Management
Slack
Apply
$184k – $288k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
AI/ML
Machine Learning
Apply
$108k – $184k per year • Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
AI/ML
Machine Learning
DevOps
Git
Management
Linear
Confluence
Jira
Agile
Apply
$224k – $357k per year • In office • Full-Time • 10+ years exp • PhD • Santa Clara
AI/ML
AI Agents
Machine Learning
Apply
$148k – $224k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
AI/ML
Reinforcement Learning
Synthetic Data
Physical AI
Game Dev
NVIDIA PhysX
Robotics
Isaac Sim
Isaac Lab
Sim-to-Real
Imitation Learning
Reinforcement Learning
Digital Twin
Apply
Technical Writer 1 day ago
≈ $120k – $250k per year (Estimated) • Equity • In office • Full-Time • Santa Clara
Design
SolidWorks
Management
Microsoft Office
Apply
≈ $199k – $359k per year (Estimated) • In office • Bachelor's Degree • Santa Clara
AI/ML
AI Agents
Apply
≈ $47k – $81k per year (Estimated) • In office • Internship • Santa Clara
Management
Slack
Google Sheets
Apply
$213k – $288k per year • Equity • In office • Full-Time • 7+ years exp • PhD • Santa Clara
AI/ML
Machine Learning
DevOps
AWS
SLI/SLO/SLA
Apply
≈ $121k – $239k per year (Estimated) • In office • Santa Clara
Apply
See all jobs
This is one of many
931,896 more open roles from verified company boards, updated every day.