654,521open jobs
38,054companies
91,731added this week
Browse all
Salary
$168k – $270k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is driving AI and high-performance computing forward. DGX Cloud aims to deliver a fully managed AI platform on major cloud providers, optimizing AI workloads using high-performance NVIDIA infrastructure. Work with NVIDIA's DGX Cloud team as a Senior Site Reliability Engineer to maintain high-performance DGX Cloud clusters for AI researchers and enterprise clients worldwide.

What makes this opportunity outstanding is that you will be at the forefront of technology, working with innovative AI and cloud computing solutions. You will have the chance to contribute to a world-class team that is determined to push the boundaries of innovation and flawlessly implement ambitious projects!

What you’ll be doing:

  • Build, implement and support operational and reliability aspects of large-scale Kubernetes clusters with focus on performance at scale, real-time monitoring, logging, and alerting.

  • Define SLOs/SLIs, monitor error allowances, and streamline reporting.

  • Support services before they launch through system creation consulting, developing software tools, platforms and frameworks, capacity management, and launch reviews.

  • Maintain services once they are live by measuring and supervising availability, latency, and overall system health.

  • Operate and optimize GPU workloads across AWS, GCP, Azure, OCI, and private clouds.

  • Scale systems sustainably through mechanisms like automation and evolve systems by pushing for changes that improve reliability and velocity.

  • Lead triage and root-cause analysis of high-severity incidents.

  • Practice balanced incident response and blameless postmortems.

  • Participate in on-call rotation to support production services.

What we need to see:

  • BS in Computer Science or related technical field, or equivalent experience.

  • 8+ years of experience operating production services.

  • Expert-level knowledge of Kubernetes administration, containerization, and microservices architecture.

  • Experience with infrastructure automation tools (e.g., Terraform, Ansible, Chef, Puppet).

  • Proficiency in at least one high-level programming language (e.g., Python, Go).

  • In-depth knowledge of Linux operating systems, networking fundamentals (TCP/IP), and cloud security standards.

  • Solid grasp of SRE principles, such as SLOs, SLIs, error budgets, and incident management.

  • Experience building and operating comprehensive observability stacks (monitoring, logging, tracing) using tools like OpenTelemetry, Prometheus, Grafana, ELK Stack, Lightstep, Splunk, etc.

Ways to stand out from the crowd:

  • Operating GPU-accelerated clusters with KubeVirt in production.

  • Applying generative-AI techniques to reduce operational toil.

  • Experience with workflow orchestration platforms such as Temporal, Cadence, Airflow, Argo Workflows, or Step Functions.

  • Experience operating and resolving problems in production AI inference workloads across the model-to-GPU stack, including vLLM, SGLang, PyTorch, TensorRT-LLM, NVIDIA Dynamo, CUDA, NCCL, and GPU performance analysis.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 19, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
654,521 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$224k – $357k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
Python
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara • Westford • Austin • Durham
Python
AI/ML
InfiniBand
DevOps
Cilium
Envoy
SLURM
Kubernetes
Cloudflare
HPC
Cybersecurity
Calico
Apply
$109k – $217k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Reston
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Databases
MongoDB
Frontend
Vue.js
Angular
Bootstrap
React.js
DevOps
Ansible
Azure
CI/CD
AWX
Management
Scrum
Kanban
QA
Swagger
Apply
$140k – $200k per year • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Reston
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Databases
MongoDB
Frontend
Vue.js
Angular
Bootstrap
React.js
DevOps
Ansible
Azure
CI/CD
AWX
Management
Scrum
Kanban
QA
Swagger
Apply
$24k – $27k per year • Remote/Hybrid • 3+ years exp • Novosibirsk
Python
Java
SQL
Bash
Databases
PostgreSQL
AI/ML
AI Agents
NLP
LLM
DevOps
Rest API
Zabbix
Prometheus
CI/CD
Git
Docker
Nginx
Grafana
Analytics
ETL/ELT
Apply
$69k – $158k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Taipei • Hsinchu
Apply
$83k – $207k per year (Estimated) • In office • Full-Time • 3+ years exp • Reading
Apply
$184k – $288k per year • In office • Full-Time • PhD • Santa Clara
Apply
$224k – $357k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
Python
Apply
$84k – $199k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • United Kingdom
AI/ML
Edge AI
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
C++
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
Python
AI/ML
LLM
RAG
Agentic Workflows
DevOps
Git
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
CUDA Toolkit
JAX
PyTorch
RAG
CUDA
DevOps
Docker
Kubernetes
HPC
Apply
$224k – $357k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • New York • Santa Clara
AI/ML
CUDA Toolkit
CUDA
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara
Python
JavaScript
Node JS
Databases
MySQL
ElasticSearch
Frontend
React.js
DevOps
CI/CD
AWS
Amazon S3
Apply
See all jobs
This is one of many
654,521 more open roles from verified company boards, updated every day.