NVIDIA DGX Cloudis an AI Factory designed to power the next generation of AI and industrial-scale breakthroughs. As the Distinguished Engineer for Security Architecture, within our Security Engineering organization, you will set the security design bar for an AI factory of hundreds of thousands of GPUs, and then build against it alongside the teams. This is the founding architecture seat in a new organization. Security Engineering is a new organization at DGX Cloud, accountable for the security outcome of the platform, and this is the architecture function inside it. You will define the security design standard for DGX Cloud, a bar that sits above the company floor, and hold it from inside the teams doing the building.
Security here is fleet horizontal and stack vertical, so your scope runs from the hardware root of trust and the hardened baseline, through tenancy and GPU workload isolation, to the services and APIs built on top, across every DGX Cloud engineering organization. A small team of Principal Engineers will report to you and hold the bar at domain depth. This is still a hands-on seat, and you stay in the design with them. You will also serve as DGX Cloud's technical interface into NVIDIA's central security organization. There is no architecture review board here and no approval queue; the bar holds because the strongest security engineers in the room helped set it and helped ship it.
What You Will Be Doing:
Set the DGX Cloud Security Bar:Own the security design standard across DGX Cloud (tenancy, GPU workloads, identity, supply chain, and isolation) and make it concrete. Reference architectures, golden paths, and requirements engineers can actually build against, not a policy library.
Hold the Bar by Building:Embed with engineering teams on real work: join the design, learn the code, help ship the thing rather than grade it afterward. The posture is not "you did this wrong." It is "here are the considerations we need to meet, I will help, let's go to work."
Architect the Security Platform:Drive the multi-year architecture for the common security infrastructure the fleet runs on: hardened baselines and patch SLOs, deploy-time policy and admission control, workload identity, supply-chain provenance through signing and attestation, and detection and response tuned to GPU workloads.
Engineer Classes of Risk Out of Existence:Finding a vulnerability is 4% of the work; building the system where it cannot happen again is the other 96%. You will attack recurring risk at its architectural root, including offensively against our own designs, and feed what you find back into the paved roads, the baselines, and the detections.
Grow the Principal Engineering Bench:Lead a small team of Principal Engineers who report to you and hold the bar across domains. Set their technical direction, hire into the gaps, and grow each of them into deeper scope than they arrived with.
Multi-Functional Collaboration:Serve as DGX Cloud's primary technical connection into NVIDIA's central security organization, and partner across platform, production engineering, and network security teams. We operate horizontally across vertically organized teams, with influence rather than authority.
Communicate Risk Upward:When something is going to ship below the bar, write down the gap and what closing it would take, and put it in front of the executives who hold the risk pen. We do not block releases ourselves and we do not quietly absorb risk. The bar holds because gaps become visible and owned.
What We Need to See:
Security Engineering:Experience (typically 20+ years) across software engineering, infrastructure, and security, with recent depth securing large-scale distributed or cloud platforms. You build systemic solutions rather than performing manual operations or "tool administration."
Architecture at Fleet Scale:A proven record owning the security architecture of a large platform end to end, across multi-tenant isolation, identity and access, policy enforcement, supply chain, and detection. You saw that architecture get built and operated, not shelved.
Production-Grade Coding:Deep enough in the code to stay credible and hands-on. You have written, shipped, and operated production services at scale yourself, and you can read an unfamiliar codebase closely enough to see where its security assumptions break.
Emerging Technology:You form a view early on what is changing under this platform, and you can tell a real shift from a trend. Today that means AI-generated code arriving faster than people can review it, AI-assisted security analysis, and agents acting inside the systems you defend. Tomorrow it will mean something else.
Threat Modeling:The depth to threat model a complex distributed system (the substrate, the orchestration layer, the tenancy boundary) and to tell an architectural gap from a bug.
Distributed Systems and Linux Depth:Cloud-native architecture, container orchestration (Kubernetes), Linux systems security including kernel-level primitives, and the security properties of high-throughput multi-tenant environments.
Technical Leadership and Influence:Build consensus and organizational alignment across engineering leadership and the most senior executives. You can turn a fast-moving risk picture into something an executive can act on, declare the unknowns honestly, and make trade-offs legible to the people funding them.
Leading Senior Engineers:Experience (typically 7+ years) leading and growing senior technical talent, including engineers deeper than you in their own domain. A small team of Principal Engineers reports to you here, so hiring, coaching, and setting direction are part of the job.
Foundation:Bachelor's degree in Computer Science, Engineering, or a related technical field (and we embraceequivalent experience 100%).
BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field - or equivalent experience with a minimum of 18+ years of practical engineering experience
Ways To Stand Out from the Crowd:
HPC/AI Security:Experience securing high-performance computing environments, RDMA-based networks, large GPU fleets, or GPU-specific security challenges.
Building the Function, Not Joining It:You have stood up a security architecture practice where none existed, including the bar, the reference architectures, the working agreements, and the credibility that made teams want it there.
Hardware Root of Trust and Workload Identity:Depth in attestation, TPM/HSM integration, confidential computing, or workload identity frameworks deployed at scale.
Open Source and Industry Impact:Notable contributions to security-focused open source, standards work, or engineering-focused security research. How have you moved the bar for the industry and not only for your employer?
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars. NVIDIA is looking for great people like you to help us accelerate the next wave of artificial intelligence.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 320,000 USD - 488,750 USD.You will also be eligible for equity and benefits.
Applications for this job will be accepted at least until September 25, 2026.This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
