410,350open jobs
14,253companies
73,765added this week
Browse all
Salary
$168k – $270k per year
Location
In office (Santa Clara)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables unique creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. NVIDIA is looking for phenomenal people like you to help us accelerate the next wave of artificial intelligence.

Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a crucial role in designing, implementing, and optimizing on-prem High-Performance Computing (HPC) storage solutions while harnessing the power of cloud computing. You will be responsible for crafting and deploying distributed storage solutions, build automation tools, and ensuring the efficient operations of our growing IT ecosystem. You will collaborate closely with engineering teams to align infrastructure with their evolving needs, document best practices, and contribute to the success of ground breaking projects.

What You'll Be Doing:

  • Design, implement an on-prem HPC infrastructure supplemented with cloud computing to support the growing IT needs of NVIDIA.

  • Design and implement scalable and efficient Storage solutions tailored for data-intensive applications, optimizing performance and cost-effectiveness.

  • Develop tooling to automate deployment and management of large-scale infrastructure environments, to automate operational monitoring and alerting, and to enable self-service consumption of resources.

  • Document the general procedures and practices, perform technology evaluations, related to distributed file systems.

  • Collaborate across teams to better understand developers' workflows and gather their infrastructure requirements.

  • Influence and guide methodologies for building, testing, and deploying applications to ensure optimal performance and resource utilization.

What we need see:

  • BS in Computer Science (or equivalent experience) with 8+ years of relevant experience, MS with 5+ years of experience or Ph.D. with 3 years of experience.

  • 8+ years of experience crafting technology solutions and resolving performance bottlenecks for HPC applications.

  • Design, deployment and management of Enterprise NAS solutions like NetApp, Pure Storage and S3 based storage such as cloudian MinIO. Design, deployment, and management of Enterprise NAS solutions like NetApp, Pure Storage, and S3-based storage such as Cloudian MinIO.

  • Experience with one or more parallel or distributed filesystems such as Lustre, GPFS.

  • Python/Bash/Golang programming/scripting experience.

  • Strong Experience operating services in any of the leading Cloud environment [ AWS, Azure or GCP].

  • Experience with multiple monitoring stacks such as Prometheus+Grafana, Elasticsearch+Kibana, Splunk, Zabbix, etc. Familiarity with newer and emerging monitoring products.

  • Excellent communication and collaboration skills.

Ways To Stand Out Of The Crowd:

  • Background with RDMA (InfiniBand or RoCE) fabrics.

  • Prior Experience with HPC cluster management tools such as Slurm, PBS, LSF, etc.

  • Experience with containerization technologies, such as Docker, Mesosphere DCOS, Kubernetes (k8s).

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative and autonomous, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 270,250 USD for Level 4, and 208,000 USD - 333,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 5, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
410,350 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$40k – $113k per year (Estimated) • Remote • Bachelor's Degree
PHP
PHP
Composer
Laravel
PHPUnit
Symfony
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon S3
AWS
AWS Lambda
CI/CD
CloudFormation
Docker
Git
IAM
Terraform
Kubernetes
Cybersecurity
Checkmarx
Dependabot
OWASP Top 10
Snyk
Veracode
Apply
DevOps 4 days ago
$73k – $154k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Montreal
Java
SQL
TypeScript
JavaScript
Java
Spring Boot
Databases
Azure SQL Database
Frontend
Angular
DevOps
ArgoCD
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
Dynatrace
Git
GitHub
GitHub Actions
GitOps
Grafana
Helm
IAM
Kubernetes
OpenShift
Prometheus
Splunk
QA
Gatling
Apply
$19k – $50k per year (Estimated) • Remote • Full-Time • 4+ years exp • Bachelor's Degree • India
Bash
PowerShell
Python
DevOps
AWS
Azure
GCP
Cybersecurity
MITRE ATT&CK
Wireshark
Apply
SOC Lead 4 days ago
$27k – $66k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • India
PowerShell
Python
DevOps
AWS
Azure
GCP
Incident Management
Cybersecurity
Carbon Black
Crowdstrike
Apply
$206k – $286k per year • In office • Internship • 5+ years exp • High School Diploma • Seattle
Java
Python
Ruby
DevOps
AWS
Azure
GCP
Management
Stripe
Apply
$34k – $82k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Apply
$224k – $357k per year • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
JAX
PyTorch
Reinforcement Learning
TensorFlow
Robotics
Digital Twin
Isaac Sim
MuJoCo
Perception
Reinforcement Learning
Sim-to-Real
Apply
$200k – $322k per year • Remote • Full-Time • 12+ years exp • PhD • United States
Python
DevOps
Amazon EKS
gRPC
Kubernetes
AWS
HPC
Apply
$64k – $227k per year (Estimated) • In office • Full-Time • 2+ years exp • PhD • Israel
Python
DevOps
CI/CD
Apply
$191k – $397k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
AI/ML
CUDA
CUDA Toolkit
NCCL
DevOps
Kubernetes
Apply
$221k – $387k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
DevOps
Incident Management
Management
ServiceNow
Apply
$56k – $94k per year • In office • Contractor • Santa Clara
Apply
$64k – $74k per year • In office • Contractor • Santa Clara
Apply
$56k per year • In office • Contractor • Santa Clara
Apply
$84k – $94k per year • In office • Contractor • Santa Clara
Apply
See all jobs
This is one of many
410,350 more open roles from verified company boards, updated every day.