401,286open jobs
13,975companies
77,769added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, Austin, Hillsboro)
Seniority
Architect · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are now looking for a Senior Hardware Architect for our Tegra System-on-Chips (SoC) focused on Reliability, Availability, and Serviceability (RAS). Do you want to be part of the Artificial Intelligence (AI) revolution and help define resilient computing platforms for datacenters, autonomous vehicles, edge systems, and other high-reliability applications? We are looking for an exceptional SoC architect to help define, drive, and deliver RAS hardware architecture across advanced CPUs and SoCs, from early architectural concepts through design implementation, verification, validation, and production readiness.

This position offers the opportunity to have real impact in a dynamic, technology-focused company developing state-of-the-art processor and system architectures at the forefront of machine learning, autonomous vehicles, high-performance computing, and edge computing. You will work with world-class systems architects, RAS experts, design teams, verification teams, validation teams, firmware teams, and software partners to define end-to-end hardware RAS features that improve system resiliency, observability, debuggability, error containment, recovery, and serviceability. Space and radiation-aware design are important areas of interest for this role, including understanding how radiation effects can influence SoC reliability, but the primary focus is broad SoC RAS architecture and driving features successfully through the product development flow.

What you’ll be doing:

  • Define and drive SoC-level RAS hardware architecture across CPUs, interconnects, memory systems, IOs, safety islands, firmware interfaces, and platform-level components.

  • Own RAS features from concept through architecture specification, micro-architecture alignment, RTL implementation support, design verification, silicon validation, debug, and production readiness.

  • Develop architectural requirements for fault detection, correction, containment, isolation, telemetry, error reporting, recovery, graceful degradation, serviceability, and diagnostic observability.

  • Work closely with design, verification, validation, firmware, software, and platform teams to ensure RAS features are implementable, verifiable, debuggable, and aligned with system-level requirements.

  • Understand the broader SoC architecture and identify how RAS mechanisms interact with performance, power, reset flows, clocks, memory hierarchy, interconnect behavior, firmware-visible controls, and platform software.

  • Create hardware specifications, architectural requirements, error-handling flows, design guidance, test plans, and architectural models in SystemC, C/C++, Python, or other relevant modeling environments where applicable.

  • Plan and review verification and validation strategies for RAS mechanisms, including error injection, recovery validation, coverage analysis, resiliency modeling, and cross-functional architecture reviews.

  • Assist in failure analysis and silicon debug for lab, post-silicon, production, and field findings; develop diagnostic screens and localization methods for latent, intermittent, and environment-sensitive failures.

  • Apply RAS architecture principles to high-reliability deployment environments, including space-aware and radiation-aware use cases where single-event effects, memory corruption, logic corruption, or cumulative radiation exposure may impact system reliability.

  • Follow industry standards and best practices related to RAS, functional safety, semiconductor reliability, debuggability, verification, validation, and silicon testing.

  • Patent novel hardware architecture techniques that improve system resiliency, observability, serviceability, and recovery.

What we need to see:

  • MS or PhD degree in computer engineering, electrical engineering, or equivalent experience.

  • At least 8+ years of SoC architecture, design, verification, reliability, silicon validation, or related hardware development experience.

  • Strong understanding of Reliability, Availability, and Serviceability (RAS) in the SoC context, including fault detection, correction, containment, telemetry, recovery, degradation modes, debug visibility, and serviceability mechanisms.

  • Experience defining and driving hardware architecture features through the full development lifecycle, including architecture definition, design implementation, verification planning, validation, debug, and production readiness.

  • Strong understanding of overall SoC architecture and the ability to reason across micro-architecture, full-chip integration, firmware interfaces, software-visible behavior, platform flows, and customer use cases.

  • Meaningful industry expertise in one or more SoC architecture areas such as RAS, safety, debug, clocks, resets, interconnects, memory controllers, IO technologies, platform integration, firmware-visible error handling, or diagnostic infrastructure.

  • Hands-on experience with design verification, silicon validation, fault injection, coverage analysis, resiliency modeling, diagnostic development, or reliability validation methodology.

  • Familiarity with radiation effects, space operation, or other high-reliability deployment environments is strongly valued, including understanding how hardware architecture can mitigate single-event effects and related reliability risks.

  • Excellent analytical, written, and verbal interpersonal skills with the ability to work effectively across architecture, design, verification, firmware, software, validation, and customer-facing teams.

Ways to stand out from the crowd:

  • Demonstrated history of architecting and delivering complex SoC RAS features across design, verification, validation, and production phases.

  • Deep familiarity with architectural resiliency techniques such as ECC, parity, redundancy, replay, checkpoint/restart, scrubbing, isolation, containment, telemetry, error logging, recovery flows, and graceful degradation.

  • Experience with cross-functional debug of hardware failures in simulation, emulation, post-silicon validation, production, customer deployments, or other high-reliability systems.

  • Familiarity with Design for Debug, Design for Test, Design for Reliability, silicon observability, fault-injection methodology, and coverage-driven validation flows.

  • Experience with radiation effects analysis, soft-error-rate analysis, radiation test campaigns, accelerated stress testing, heavy-ion or proton testing, or space qualification methodology.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us. If you are creative, autonomous, and motivated to define and deliver resilient SoC architectures that improve reliability, debuggability, and serviceability across demanding applications, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 6, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
401,286 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$130k – $180k per year • Remote • 10+ years exp • Bachelor's Degree
C++
C
C++
LLVM
PyTorch C++
TensorFlow C++
C
MPI
AI/ML
CUDA
CUDA Toolkit
CUTLASS
DeepSpeed
JAX
MLIR
NCCL
PyTorch
ROCm
TensorFlow
TensorRT
Triton
vLLM
DevOps
AWS
Azure
GCP
HPC
Apply
$100k – $170k per year • Remote • 12+ years exp • Bachelor's Degree
Python
DevOps
IAM
Terraform
Cybersecurity
CIS Benchmarks
HIPAA
ISO 27001
PCI DSS
SOC 2
Zero Trust
Apply
$100k – $113k per year • Remote • 7+ years exp • Bachelor's Degree
Python
DevOps
IAM
Terraform
Cybersecurity
CIS Benchmarks
HIPAA
ISO 27001
PCI DSS
SOC 2
Zero Trust
Apply
$86k – $103k per year • Remote • 8+ years exp • Bachelor's Degree
Bash
Go
Python
DevOps
Ansible
ArgoCD
AWS
Azure
CI/CD
GCP
GitOps
Helm
Istio
Jenkins
Kubernetes
Linkerd
OpenShift
Red Hat
Service Mesh
Tekton
Terraform
Cybersecurity
HIPAA
PCI DSS
SOC 2
Apply
$100k – $150k per year • Remote • 6+ years exp • Bachelor's Degree
Python
Databases
HeatWave
MySQL
Oracle
DevOps
AWS
Azure
CI/CD
FinOps
GCP
GitHub
GitHub Actions
IAM
Jenkins
Kubernetes
Terraform
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
Agentic Workflows
AI Agents
Function Calling
LangChain
LlamaIndex
LLM
OpenAI
Tool Use
DevOps
CI/CD
Git
Apply
In office • Internship • PhD • Shanghai
C++
Python
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenCL
DevOps
HPC
Apply
In office • Internship • Master's Degree • Shanghai
C#
C++
Python
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
AI Agents
Apply
In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$171k – $324k per year (Estimated) • Equity • In office • 15+ years exp • Master's Degree • Santa Clara
AI/ML
AI Agents
DevOps
AWS
Azure
GCP
Cybersecurity
Zero Trust
Marketing
Instagram
LinkedIn
Apply
$100k – $137k per year • Equity • In office • Full-Time • 3+ years exp • Master's Degree • Santa Clara
Chips/EDA
Cadence Allegro
OrCAD
Apply
$114k – $228k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Marketing
X (Twitter)
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
See all jobs
This is one of many
401,286 more open roles from verified company boards, updated every day.