679,351open jobs
39,352companies
99,642added this week
Browse all
Salary
$48k – $108k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Staff · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking a Senior Staff SRE to build and operate reliable, scalable compute platforms that support global engineering workloads. This role spans Kubernetes, KubeVirt, bare-metal infrastructure, automation, observability, and AI-enabled operations. Join a team that solves complex infrastructure challenges, builds durable automation, and improves the reliability and operational experience of critical compute services.

What you’ll be doing:

  • Build, operate, and improve large-scale Kubernetes, KubeVirt, Linux, container, and bare-metal compute platforms, with a focus on performance, capacity, reliability, and operational scale.

  • Lead bare-metal provisioning and lifecycle management in data centers, including PXE boot, DHCP, DNS, OS provisioning, hardware validation, and fleet automation.

  • Develop automation, self-service capabilities, and observability solutions using APIs, Python or Go, Infrastructure as Code, configuration management, metrics, logs, traces, and service-health data.

  • Define and operate SLOs, SLIs, error budgets, alerting, and incident-response practices; lead complex incident investigations, corrective actions, and blameless postmortems.

  • Partner with infrastructure, security, hardware, data-center, and application teams to deliver global platform initiatives, and participate in an on-call rotation.

What we need to see:

  • BS in Computer Science, Engineering, a related technical field, or equivalent experience, plus 10+ years operating production infrastructure or platform services.

  • Strong expertise in Kubernetes administration, KubeVirt, Docker, containerization, microservices, Linux systems, and resolving distributed-system challenges.

  • Experience deploying and operating bare-metal infrastructure in a data-center environment, including provisioning, networking, operating-system lifecycle management, and hardware automation.

  • Proficiency in Python, Go, or a comparable programming language, with experience building RESTful services and integrating infrastructure APIs.

  • Experience with Infrastructure as Code and automation tools such as Terraform, Ansible, Chef, or Puppet, along with a solid understanding of TCP/IP networking and infrastructure security.

  • Strong SRE and observability experience, including SLIs, SLOs, error budgets, incident management, monitoring, logging, tracing, and tools such as OpenTelemetry, Prometheus, Grafana, ELK Stack, or Splunk.

  • Clear written and interpersonal communication skills, with a record of delivering practical, scalable solutions to complex technical problems.

Ways to stand out from the crowd:

  • Experience operating HPC, AI, GPU-accelerated, or general-purpose bare-metal compute infrastructure, including GPU-enabled Kubernetes or KubeVirt clusters.

  • Expertise with VMware vSphere, Red Hat OpenShift, KVM, Firecracker, OpenStack, or Nutanix AHV.

  • Experience applying generative AI or agentic workflows to improve infrastructure diagnostics, reduce operational toil, and accelerate incident resolution.

  • Experience building secure, integrated operational platforms using APIs, RBAC, service accounts, secrets management, audit controls, workflow orchestration, and infrastructure or incident-management systems.

  • Demonstrated delivery of complex, high-impact infrastructure projects.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
679,351 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
In office • 6+ years exp • PhD
Python
Go
DevOps
Terraform
GCP
Istio
OpenTelemetry
Datadog
Linkerd
Prometheus
Azure
AWS
Kubernetes
Grafana
Service Mesh
Management
Agile
Scrum
Apply
$15k – $33k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Manila
DevOps
Terraform
Azure
CI/CD
Docker
Kubernetes
Azure AKS
Apply
$9.5k – $22k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Manila
DevOps
Terraform
Azure
CI/CD
Git
Docker
Kubernetes
Azure AKS
Apply
Manager, Market Risk 2 hours ago
$139k – $231k per year • In office • Full-Time • 8+ years exp • Master's Degree • New York • Chicago
Python
SQL
Apply
Data Engineer II 2 hours ago
$21k – $49k per year (Estimated) • In office • Mumbai
Python
SQL
Databases
Snowflake
AI/ML
Copilot
Claude
Analytics
Tableau
ETL/ELT
Apply
$88k – $138k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Apply
$184k – $299k per year • In office • Full-Time • 12+ years exp • Master's Degree • Santa Clara
Apply
$152k – $242k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Santa Clara • Austin • Redmond
C++
C++
TensorFlow C++
PyTorch C++
LLVM
AI/ML
CUDA Toolkit
JAX
OpenCL
TensorFlow
PyTorch
CUDA
Triton
OpenAI
MLIR
Apache TVM
XLA
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Redmond • Santa Clara
C++
C++
LLVM
AI/ML
CUDA Toolkit
CUDA
MLIR
DevOps
HPC
Apply
$95k – $219k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
Lua
DevOps
SLURM
HPC
Apply
$26k – $55k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • Bengaluru
Apply
$35k – $72k per year (Estimated) • In office • Full-Time • Bengaluru • Hyderabad
Apply
$20k – $43k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Hyderabad • Bengaluru
Python
JavaScript
Java
Node JS
DevOps
Terraform
Ansible
Azure
CI/CD
AWS
Kubernetes
Management
Agile
Scrum
Apply
$28k – $58k per year (Estimated) • In office • Full-Time • 14+ years exp • Bengaluru
JavaScript
TypeScript
Frontend
Vue.js
Redux
Angular
React.js
Sass
NgRx
Mobile
State Management
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Design
Figma
Management
Agile
Apply
$31k – $78k per year (Estimated) • In office • 4+ years exp • Bengaluru
Java
Kotlin
Dart
Swift
Mobile
SwiftUI
Flutter
Apply
See all jobs
This is one of many
679,351 more open roles from verified company boards, updated every day.