691,154open jobs
40,551companies
97,801added this week
Browse all
Location
In office (Courbevoie, United Kingdom, Spain, Germany)
Seniority
Architect
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is looking for a Senior Cloud Infrastructure and DevOps Solutions Architect to join its NVIDIA Infrastructure Specialist Team. Academic and commercial organizations around the world are using NVIDIA products to redefine deep learning and data analytics, and to power next-generation data centers. Join the team building and advising on many of the largest and fastest AI/HPC systems in the world!

We are looking for someone who combines deep technical expertise with strong consulting and communication skills. This role will engage directly with customers, partners, and cross-functional teams to assess, architect, and guide the implementation of large-scale infrastructure projects. The scope spans system architecture, Kubernetes-based platforms, and automation-serving as both a trusted advisor and a hands-on technical leader. You will sit at the centre of NVIDIA's Cloud Partner (NCP) operating model, covering the full Day 1 to Day 2 lifecycle: taking a GPU cluster from hardware handover, through full-solution validation, to a production-stable platform running at maximum goodput. NCP estates are open-source-first and heterogeneous-upstream Kubernetes, KubeVirt, Slurm, Prometheus/Grafana, Cumulus/SONiC and a long tail of ISV software-so this role is deliberately tool-agnostic: you will meet each partner on the stack they actually run rather than on a single proprietary product.

What You’ll Be Doing:

  • Own full-solution validation on the partner software stack-the layer above hardware validation-including cluster-wide stability testing, real training-workload acceptance, and multi-day, multi-rack burn-in against agreed MTBI and goodput targets.

  • Minimise the time from cluster handover to first production workload, working across hardware bring-up, managed-service intake and the partner's own operations teams to remove duplicated validation and handover friction.

  • Own Day 2 production stability at fleet scale: monitoring, logging and workload orchestration, fault detection and remediation, preventive maintenance, and proactive firmware and field-notice rollout campaigns.

  • Assess customer environments and operate heterogeneous open platforms-upstream Kubernetes, KubeVirt, Slurm and GPU-aware schedulers-integrated with enterprise-grade networking and storage, and enable third-party ISV workloads on top of them.

  • Provide consultative guidance and hands-on troubleshooting across the full stack-bare metal, operating system, software stack, container platform, networking and storage-and support R&D, POCs and POVs validating new features, architectures and upgrade approaches.

  • Act as the technical leader for assigned accounts: run structured knowledge transfer and enablement, and produce runbooks, onboarding materials and best-practice guides so partner teams can operate advanced configurations independently.

What We Need to See:

  • BS/MS/PhD in Computer Science, Electrical/Computer Engineering, Physics, Mathematics, or related fields, or equivalent experience.

  • 8+ years in managing scalable cloud environments and automation engineering roles.

  • Cloud, HPC & GPU Expertise: Proven understanding of networking fundamentals and data centre architectures, with hands-on experience managing HPC/AI clusters and NVIDIA GPU-accelerated infrastructure-deployment, driver and CUDA toolkit management, optimisation, workload profiling and troubleshooting across CPUs, GPUs and high-speed interconnects.

  • Kubernetes & AI/ML Workloads: Extensive background with Kubernetes for container orchestration, resource scheduling and scaling in GPU-accelerated and HPC environments, including scheduler internals, batch schedulers such as Slurm, and mixed bare-metal/virtualised (e.g. KubeVirt) multi-tenant estates.

  • Linux & Storage Systems: Deep knowledge of Linux (RedHat, Ubuntu), OS-level security, and protocols. Experience with storage solutions such as Lustre, GPFS, ZFS, XFS, and emerging Kubernetes storage technologies.

  • Automation, GitOps & Observability: Proficiency in Python and Bash scripting, configuration management and Infrastructure-as-Code tools (e.g. Ansible, Terraform), GitOps-based cluster lifecycle and upgrade management for large fleets, and observability stacks (Grafana, Loki, Prometheus) for monitoring, logging and building fault-tolerant systems.

  • Fleet Reliability & Customer Engagement: Demonstrated ability to measure and improve MTBI and job goodput on large GPU clusters-fault detection, drain and remediation workflows, SLO/error-budget definition and post-incident review-combined with a strong consultative background leading architectural reviews and presenting to executive stakeholders.

Ways to Stand Out from the Crowd:

  • Knowledge of CI/CD pipelines and container-based microservices architectures for software deployment and automation.

  • Experience with the NVIDIA GPU and Network Operators for automated GPU and network resource lifecycle management in Kubernetes, and with NVIDIA Base Command Manager (BCM) for provisioning, managing and monitoring GPU clusters at scale.

  • Familiarity with GPU health and fleet telemetry tooling-DCGM and XID diagnostics, node-level health agents, and fleet-wide reliability intelligence.

  • Expertise in AI-native scheduling and inference frameworks on Kubernetes (e.g. KAI, Grove, Dynamo, NVIDIA Cloud Functions).

  • Background with RDMA-based fabrics (InfiniBand or RoCE) in HPC or AI environments. Exposure to Cumulus Linux, SONiC or Spectrum-X fabrics, DPU/DOCA infrastructure services, and NVLink/NVSwitch partition operations (NMX-C / NMX-M) on NVL72-class systems is a strong plus.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
691,154 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Courbevoie
$72k – $185k per year (Estimated) • Remote/Hybrid • Internship • High School Diploma • Chicago • Schaumburg
Python
Java
AI/ML
LLM
DevOps
Splunk
GCP
Azure
AWS
Linux
Windows
TCP/IP
Cybersecurity
Wireshark
Nmap
Tcpdump
Wazuh
Volatility
Suricata
Autopsy
Velociraptor
osquery
OSSEC
FTK
SIEM
Apply
$68k – $157k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Atlanta
Python
JavaScript
TypeScript
SQL
Python
pySpark
Databases
Snowflake
ElasticSearch
AI/ML
Hadoop
Spark
Amazon SageMaker
DevOps
AWS
Kubernetes
Apply
$132k – $207k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Python
AI/ML
InfiniBand
DevOps
HPC
Linux
Apply
$124k – $196k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin
Python
Apply
$224k – $357k per year • Remote • Full-Time • 12+ years exp • Bachelor's Degree • United States
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorFlow
PyTorch
Synthetic Data
CUDA
Machine Learning
Apply
OEM Sales Director 2 hours ago
$296k – $449k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
Apply
$132k – $207k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Python
AI/ML
InfiniBand
DevOps
HPC
Linux
Apply
$124k – $196k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Austin
Python
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Santa Clara
Management
Outlook
Microsoft Office
Apply
$124k – $227k per year (Estimated) • Remote • Full-Time • 8+ years exp • Australia
Go
AI/ML
Edge AI
Apply
$92k – $263k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Zurich • Courbevoie
Python
DevOps
Linux
Apply
$24k – $55k per year (Estimated) • In office • Internship • Courbevoie
Apply
$35k – $80k per year (Estimated) • In office • Full-Time • 5+ years exp • Courbevoie
Design
Figma
Apply
$41k – $74k per year (Estimated) • Remote/Hybrid • Full-Time • PhD • Courbevoie
Apply
$32k – $59k per year (Estimated) • Remote • Full-Time • Courbevoie
JavaScript
Java
TypeScript
SQL
Java
Spring Boot
Hibernate
Frontend
Angular
DevOps
Jenkins
Management
Confluence
Apply
See all jobs
This is one of many
691,154 more open roles from verified company boards, updated every day.