402,911open jobs
14,044companies
78,108added this week
Browse all
Salary
$116k – $184k per year
Location
In office (Santa Clara)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most experienced and hardworking people in the world working for us. If you're creative, autonomous, and energized by deep technical problem solving, we want to hear from you!

We are looking for a Product Quality Engineer to join our Systems Product Quality team as the system-level power debug domain expert for customer returns and field failures. This role will lead technical failure analysis for NVIDIA data center systems, compute trays, and compute modules, with special focus on customer-reported field issues involving power delivery, power sequencing, and intermittent power events.

Your hands-on debug capability, structured root-cause mindset, and ability to connect board-level signals with system-level behavior will be essential to driving customer-return investigations from symptom confirmation through technical root cause and corrective action closure. Own customer-return and field-failure power investigations from symptom confirmation through containment, root cause, corrective action, and quality learning closure.

What you'll be doing:

  • Lead system-level power failure analysis for customer returns and field failures across data center systems, compute trays, and compute modules.

  • Confirm, reproduce, and isolate complex power failures such as no power, intermittent boot, unexpected shutdown, brown-out, rail droop, over-current protection, under-voltage protection, sequencing faults, hot-plug events, and margin-related failures.

  • Analyze system power architecture from AC/DC input through PSU, PDU, hot-swap, eFuse, VR, regulator, current-sense, and board-level power rails to determine the true failure boundary.

  • Use oscilloscopes, current probes, DMMs, BMC-reported voltage/current readings, system event logs, and Linux-based diagnostics to build fact-based debug conclusions.

  • Correlate field return data, customer logs, firmware behavior, board schematics, PCB layout, BOM history, and telemetry trends to identify root cause and assess risk.

  • Partner with hardware design, power design, firmware, customer quality, reliability, manufacturing, and supplier quality teams to resolve critical customer and field issues.

  • Drive containment, failure analysis, corrective and preventive actions, and defect-prevention feedback with clear ownership and closure criteria.

  • Create concise technical reports, quality updates, and executive-ready summaries that communicate failure mechanism, impact, risk, mitigation, and next steps.

What we need to see:

  • Bachelor's degree or equivalent experience in Electrical Engineering, Electronic Engineering, or a related field; Master's degree preferred.

  • 5+ years of hands-on experience in hardware debug, customer return analysis, field failure analysis, or power electronics support for complex electronic systems.

  • Strong understanding of system power delivery, DC-DC converters, multiphase VRs, regulators, power sequencing, current sharing, sense circuits, protection circuits, and high-current low-voltage rails.

  • Proven ability to debug power issues at system, board, and component level by reading schematics, PCB layouts, power trees, design specifications, and test logs.

  • Experience with Linux systems, Linux shell scripts, BMC/IPMI/Redfish-style logs or telemetry, and basic automation for data collection and debug efficiency.

  • Strong analytical and problem-solving skills, including structured troubleshooting, design of experiments, root cause analysis, statistical process control, and quality data analysis.

  • Ability to work across engineering, customer quality, supplier, and customer-facing teams while maintaining clear technical ownership and urgency.

  • Excellent written and spoken English, strong documentation habits, and the ability to explain complex debug findings to both technical and non-technical audiences.

  • High sense of responsibility, self-motivation, collaborative working style, and comfort driving ambiguous technical issues to closure.

Ways to stand out from the crowd:

  • Experience debugging high-power server or data center platforms in customer-return or field-failure analysis workflows.

  • Hands-on familiarity with PSU/PDU behavior, rack-level power distribution, power capping, power transients, or data center deployment conditions observed in field returns.

  • Experience with board-level power design, hardware verification, power integrity measurement, or design-for-debug improvements.

  • Knowledge of quality and reliability concepts, 8D problem solving, customer failure reporting, RMA/FA workflow, and supplier corrective action processes.

  • Ability to confirm, bound, and translate power-related field failures into corrective actions, debug playbooks, and prevention feedback for design, customer quality, and supplier teams.

With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, with a genuine passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 184,000 USD for Level 3, and 148,000 USD - 235,750 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 18, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
402,911 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 3 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
$171k – $324k per year (Estimated) • Equity • In office • 15+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Cybersecurity
Zero Trust
Marketing
Instagram
LinkedIn
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
See all jobs
This is one of many
402,911 more open roles from verified company boards, updated every day.