368,746open jobs
9,444companies
47,506added this week
Browse all
Location
In office (Shanghai)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

At NVIDIA, we are at the forefront of technological innovation, pushing the boundaries of AI and accelerated computing. Our team in Shanghai, China is looking for a Senior Site Reliability Engineer focused on Test Environment Management to join us. This is an opportunity to build and operate highly reliable test infrastructure, CI/CD systems, and environments that power validation of NVIDIA enterprise offerings. If you are passionate about reliability engineering, test infrastructure, and AI-scale systems, this role is for you.

What you’ll be doing

  • Design, build, operate, and continuously improve reliable, scalable test environments and automation infrastructure that support validation of NVIDIA enterprise offerings.

  • Own end-to-end CI/CD pipelines using GitLab CI, GitHub Actions, and ArgoCD (GitOps) - including pipeline design, reliability, performance, and progressive delivery of test workloads.

  • Manage Software Bills of Materials (SBOMs): generation, continuous monitoring, vulnerability correlation, policy enforcement, and integration into CI/CD and release gate..

  • Provision, scale, observe, and lifecycle-manage ephemeral and long-lived test environments (Kubernetes-based and hybrid) with strong emphasis on isolation, reproducibility, and rapid recovery.

  • Define and drive reliability practices for test systems: SLIs/SLOs/error budgets for test environments and pipelines, toil reduction, chaos/resilience testing of infrastructure, and automated remediation.

  • Collaborate closely with development and platform teams to triage environment and pipeline failures, perform root-cause analysis, verify fixes, and continuously harden test infrastructure.

  • Apply AI/ML/Agentic techniques and internal tools to accelerate environment provisioning, flaky-test detection, capacity planning, anomaly detection, and overall Quality Assurance velocity.

What we need to see

  • MS or PhD in Computer Science, related field, or equivalent experience with 8+ years of professional experience in Site Reliability Engineering, Test Environment Management, CI/CD platform engineering, or software testing infrastructure.

  • Strong proficiency with Linux, shell scripting, and Python (or equivalent automation languages).

  • Hands-on experience designing and operating CI/CD systems with GitLab CI and/or GitHub Actions, practical experience with ArgoCD (or equivalent GitOps tooling) for CD of applications and infrastructure. Solid background in containerization and orchestration (Docker, Kubernetes) and virtualization technologies.

  • Deep understanding of SRE principles: SLIs/SLOs, error budgets, incident response, postmortems, toil elimination, and reliability engineering for complex distributed systems.

  • Experience building and operating test environments (ephemeral, multi-tenant, or production-like) with focus on reliability, isolation, and rapid turnaround.

  • Strong knowledge of QA principles and how test infrastructure enables high-quality software delivery.

  • Comfort working with AI/LLM-related workloads and toolings. Excellent problem-solving skills, clear written and verbal communication, and the ability to collaborate across engineering teams.

  • Self-motivated, proactive, and passionate about learning new technology at scale.

Ways to stand out from the crowd

  • Experience operating large-scale Kubernetes platforms and GitOps workflows in production or high-stakes test environments.

  • Background in software supply-chain security, SBOM tooling ecosystems, vulnerability management, and policy enforcement (OPA/Gatekeeper, Kyverno, etc.).

  • Hands-on work with NVIDIA GPU hardware, multi-GPU environments, or accelerated computing infrastructure. Experience with parallel programming, high-performance computing, or large-scale AI model training/inference test harnesses.

  • Track record of applying AI/observability techniques to detect flaky tests, optimize environment utilization, or automate root-cause analysis.

  • Prior experience defining and driving reliability programs (error budgets, chaos engineering, capacity forecasting) for CI/CD or test platforms.

NVIDIA is widely considered one of the technology world’s most desirable employers. We have some of the most brilliant people on the planet working for us. If you’re creative, autonomous, and excited about making test infrastructure as reliable as the products it validates, we want to hear from you!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$118k – $142k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • El Segundo
C++
MATLAB
Python
SystemC
SystemVerilog
Verilog
VHDL
DevOps
CI/CD
QEMU
Chips/EDA
UVM
Apply
$129k – $232k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
C++
DevOps
CI/CD
Git
RTOS
Apply
$169k – $321k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Phoenix
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
AWS
Kong
Amazon S3
API Gateway
Cybersecurity
Zero Trust
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
Python
DevOps
CI/CD
SpaceTech
NASA cFS
Apply
$133k – $161k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Westminster
C++
C
C
U-Boot
DevOps
CI/CD
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$156k – $348k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United Kingdom
AI/ML
CUDA
CUDA Toolkit
AI Agents
NVIDIA NeMo
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
DevOps
AWS
Cybersecurity
GDPR
PCI DSS
Analytics
Power BI
Tableau
Management
Confluence
Jira
Trello
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
C++
Python
SQL
AI/ML
LangChain
LLM
Spark
Apply
In office • Full-Time • 15+ years exp • Bachelor's Degree • Shanghai
AI/ML
CUDA
CUDA Toolkit
LocalAI
DevOps
CI/CD
KVM
QEMU
Apply
In office • Full-Time • Shanghai
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.