1,244,388open jobs
71,871companies
213,835added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 5, 2026. First seen by Alion on Oct 2, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA DGX Cloud delivers AI services and endpoints for research and production workloads. We are looking for a Senior Production Engineer to build software and automation that make those services reliable, scalable, and safe to operate. The Production Engineering team works on large-scale distributed systems spanning internal and external model endpoints; regional control plane services that orchestrate workloads and route requests; and the GPU/CPU compute infrastructure where inference and agentic workloads run. Our work spans Kubernetes clusters across AWS, Azure, Google Cloud, other partner cloud environments, and on-premises deployments.

What you’ll be doing:

  • Build and operate production software, automation, and tooling for control plane services, model deployments, and inference and agentic workloads across DGX Cloud environments.
  • Improve the reliability of inference and agentic platforms and services, including NVIDIA Cloud Functions, SGLang- and vLLM-based endpoints, and inference services built with NVIDIA Dynamo, through health validation, safer rollouts, observability, and recovery.
  • Improve endpoint availability, inference routing, capacity management, and service health to maintain predictable performance as workloads and demand change.
  • Use infrastructure as code and GitOps to deploy, configure, validate, upgrade, and recover services consistently across environments.
  • Build workflows for service enablement, model releases, handoff, deprecation, and ongoing operations; replace repeatable manual work with reliable automation.
  • Define and instrument SLIs and SLOs for inference and control plane services, including availability and latency, use error budgets to guide reliability improvements, and make production health visible to partner teams.
  • Participate in on-call and incident response, troubleshoot failures across routing, model runtimes, software, and infrastructure, and turn recurring issues into automation and durable fixes.
  • Collaborate with model, platform, storage, networking, security, and GPU infrastructure teams to design and operate services safely at scale.

What we need to see:

  • 8+ years of experience building or operating production services and large-scale distributed systems, including hands-on automation.
  • Strong programming skills in Python, Go, or a comparable language, with experience developing tools for production operations.
  • Experience with infrastructure as code, configuration management, or GitOps, and with building automation for repeatable service deployments and changes.
  • Strong knowledge of Linux, Kubernetes, containers, cloud infrastructure, distributed systems, and networking fundamentals; ability to diagnose failures in production.
  • Understanding of SRE principles, including SLIs, SLOs, error budgets, incident response, and reducing operational toil.
  • Experience instrumenting services and using metrics, logs, and traces to understand system behavior and improve reliability.
  • Clear technical communication and ability to work across engineering teams.
  • BS/MS in Computer Science or equivalent experience.

Ways to stand out from the crowd

  • Familiarity with technologies such as vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, or NCCL, and with GPU performance analysis.
  • Experience building Kubernetes operators, controllers, workload orchestration services, fleet management systems, or self-healing automation.
  • Experience with Terraform, Argo CD, CI/CD, policy validation, or safe deployment and rollback systems.
  • Experience developing with AI tools and agents.
  • Background with production AI inference or agentic workloads, including debugging issues across models, runtimes, Kubernetes, and hardware.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hard-working people on the planet working for us. If you're creative, hard-working and self-motivated, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 6, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,244,388 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Santa Clara
Sr. AWS IAM Engineer 2 months ago
$70k – $170k per year • Remote (United States) • Full-Time • 5+ years exp • Santa Clara
DevOps
Terraform
CloudFormation
Azure
AWS
IAM
Cybersecurity
Least Privilege
Microsoft Entra ID
Apply
≈ $78k – $158k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
SQL
DevOps
Windows Server
Incident Management
Linux
Windows
Cybersecurity
GDPR
Active Directory
Management
ITIL
Apply
≈ $105k – $205k per year (Estimated) • In office • TS/SCI • 13+ years exp • Bachelor's Degree • Augusta
Python
JavaScript
TypeScript
PowerShell
DevOps
Ansible
Azure
CI/CD
Jenkins
Git
AWS
Linux
Management
Agile
Apply
Platform Engineer III 3 months ago
≈ $87k – $167k per year (Estimated) • Remote (United States) • Full-Time • 4+ years exp • Bachelor's Degree • United States
Python
JavaScript
DevOps
Splunk
Terraform
Ansible
Istio
New Relic
AWS
Kubernetes
Amazon EKS
Amazon EC2
IAM
API Gateway
Apply
Cloud Engineer IV 7 days ago
≈ $110k – $203k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree • United States
Python
PowerShell
Bash
DevOps
Splunk
Terraform
Ansible
GitHub Actions
Terragrunt
CloudFormation
Prometheus
GitLab CI
CI/CD
Jenkins
Git
AWS
Kubernetes
Grafana
Configuration Management
Amazon EKS
Amazon EC2
IAM
Amazon CloudWatch
Linux
Windows
Management
Jira
ServiceNow
Agile
Scrum
Apply
≈ $109k – $227k per year (Estimated) • Remote (likely Canada) • Full-Time • 7+ years exp • Bachelor's Degree
Python
Ruby
AI/ML
Model Context Protocol
RAG
Apply
≈ $93k – $218k per year (Estimated) • Remote (Canada) • Full-Time • 7+ years exp • Bachelor's Degree
JavaScript
C#
Node JS
Databases
PostgreSQL
DynamoDB
Frontend
React.js
DevOps
Terraform
AWS
GitLab
Amazon ECS
Apply
Director (R&D) 2 days ago
≈ $160k – $300k per year (Estimated) • In office • TS/SCI • 12+ years exp • Bachelor's Degree • Fairfax
Python
Java
C++
MATLAB
AI/ML
AI Agents
LLM
RAG
Machine Learning
Apply
≈ $82k – $201k per year (Estimated) • In office • Top Secret • 15+ years exp • Bachelor's Degree • Fairfax
Python
MATLAB
SpaceTech
GDAL
Apply
≈ $113k – $220k per year (Estimated) • In office • Top Secret • 15+ years exp • Bachelor's Degree • Fairfax
Python
MATLAB
AI/ML
Machine Learning
Robotics
Sensor Fusion
SpaceTech
GDAL
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara • Redmond
Python
Rust
AI/ML
Fine-tuning
Prompt Engineering
Multimodal AI
AI Agents
NLP
VLM
LLM
DevOps
CI/CD
Docker
Kubernetes
Apply
$208k – $334k per year • In office • Full-Time • 12+ years exp • Master's Degree • Santa Clara
Python
Ruby
DevOps
HPC
DNS
DHCP
VPN
BGP
MPLS
Apply
$92k – $155k per year • In office • Full-Time • Bachelor's Degree • Santa Clara
DevOps
Linux
Windows
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • PhD • Santa Clara
Python
AI/ML
Cursor
Claude
DevOps
Jenkins
Git
Gerrit
Linux
Unix
VLAN
BGP
OSPF
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
C++
DevOps
Linux
Windows
Game Dev
DirectX 12
Apply
$152k – $228k per year • In office • 8+ years exp • PhD • Santa Clara
AI/ML
AI Agents
LLM
AWS Bedrock AgentCore
LLM Guardrails
Agentic Workflows
DevOps
CI/CD
AWS
Platform Engineering
Cybersecurity
SBOM
SLSA
Quantum
IonQ
Apply
$94k – $215k per year • In office • Full-Time • 10+ years exp • Santa Clara
DevOps
Terraform
Ansible
GCP
Azure DevOps
GitHub Actions
CloudFormation
Pulumi
Azure
CI/CD
Jenkins
AWS
Kubernetes
Platform Engineering
Bicep
Amazon EKS
Google GKE
Azure AKS
FinOps
AIOps
GitLab
IAM
DNS
DHCP
Cybersecurity
ISO 27001
SOC 2
GDPR
HIPAA
Zero Trust
PKI
SIEM
Management
ITIL
Apply
≈ $41k – $71k per year (Estimated) • In office • 2+ years exp • Santa Clara
Apply
$56k – $58k per year • In office • 2+ years exp • Santa Clara
Apply
$56k – $60k per year • In office • Part-Time • Santa Clara
Apply
See all jobs
This is one of many
1,244,388 more open roles from verified company boards, updated every day.