413,737open jobs
14,642companies
73,902added this week
Browse all
Salary
$184k – $288k per year
Location
Remote (United States)
Seniority
Architect · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA's Infrastructure Specialists team is hiring a Senior Solutions Architect - AI Factory Observability & Visualization! This remote role develops full-spectrum visibility that supports the smooth functioning of HPC systems and AI factories, transforming intricate telemetry across network and compute into straightforward, actionable perspectives.

The role has a complete, end-to-end understanding of the HPC/AI system, running and interpreting microbenchmarks and workloads to confirm system readiness, then establishing the observability that maintains this state. The work involves collaborating across NVIDIA teams to help partners see, understand, and respond to HPC system and AI factory performance, from hardware to workload.

What You Will be Doing:

  • Run AI factory validation tools, microbenchmarks, and workloads provided by the team, and interpret results to assess system health and performance.

  • Gain a comprehensive understanding of the system from start to finish, including network topology, interconnects, and compute.

  • Establish what "healthy" represents across the stack - the metrics, logs, and signals that confirm a system is functioning well, and the thresholds that show it isn't.

  • Build and extend the telemetry surface across hardware, fabric, and workload, crafting how data is collected, transformed, stored, and surfaced.

  • Serve as the observability expert, investigating gaps in visibility to ensure it reflects true system behavior.

  • Develop automation (Python, Shell) for collecting, transforming, and presenting system and network data.

  • Recommend improvements to system visibility, data sources, and reporting that give teams clearer insight.

  • Collaborate with hardware, software, networking, datacenter, and product groups to ready HPC systems and AI factories for customer deployment, contributing documentation and readiness materials throughout the process.

What We Need to See:

  • Bachelor's degree or equivalent experience in Computer Science, Mathematics, Engineering, Physics, or related field.

  • 6+ years of experience managing Linux-based systems in HPC, distributed systems, or large AI/ML settings.

  • Hands-on experience with the architecture of multi-GPU and/or multi-node clusters, including networking and interconnects.

  • Solid grasp of how HPC and AI factory systems fit together end to end, from network fabric through compute.

  • Proficiency with Python and Shell/Bash for scripting, automation, and tooling.

  • Practical experience working with observability systems (e.g., Prometheus, Grafana, Loki, or similar), including building custom exporters or collectors, setting up alerts, and handling metric cardinality and retention on a large scale.

  • Experience transforming metrics, logs, and traces into clear, actionable insight for complex distributed environments.

  • Familiarity with GPU and fabric telemetry (e.g., DCGM, NVLink, InfiniBand/Ethernet fabric counters) and using it to diagnose performance regressions.

  • Strong communication skills and the ability to work effectively with cross-functional teams.

Ways to Stand Out From the Crowd:

  • Experience with AI factory or large-scale AI infrastructure build, deployment, or operations.

  • Background in HPC systems engineering, SRE, or systems analysis for GPU-accelerated environments.

  • Experience building automation and data pipelines that feed dashboards and reporting at scale.

  • Demonstrated desire to use AI to solve practical problems, improve workflows, and guide data-driven decisions.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 28, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
413,737 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Austin
$14k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Hyderabad
Bash
PowerShell
Python
SQL
DevOps
Ansible
AWS
Azure
Azure DevOps
CI/CD
CloudFormation
GCP
Git
Hyper-V
Red Hat
Terraform
VMWare
Apply
In office • Full-Time • Minsk
Bash
PowerShell
Python
DevOps
Amazon EC2
AWS
Azure
CentOS Stream
Debian
Docker
GCP
HAProxy
Hyper-V
Kubernetes
KVM
Nginx
Proxmox VE
Ubuntu
VMWare
Windows Server
Zabbix
Cybersecurity
Tcpdump
Wireshark
Apply
$17k – $45k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
Bash
Python
DevOps
Ansible
CI/CD
Debian
GitLab
GitLab CI
Prometheus
Puppet
Terraform
Ubuntu
VMWare
Windows Server
Zabbix
Docker
Management
YouTrack
Apply
ML-инженер 8 hours ago
$18k – $48k per year (Estimated) • Remote • Full-Time • Krasnoyarsk
Bash
Python
AI/ML
Keras
NLP
PyTorch
TensorFlow
Time Series Forecasting
DevOps
Docker
Git
Apply
$92k – $232k per year (Estimated) • In office • Full-Time • 7+ years exp • Singapore
Go
Databases
MySQL
Redis
DevOps
CI/CD
Datadog
Git
GitLab
Grafana
gRPC
Istio
Kibana
Kubernetes
Service Mesh
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
Agentic Workflows
AI Agents
Function Calling
LangChain
LlamaIndex
LLM
OpenAI
Tool Use
DevOps
CI/CD
Git
Apply
$14k – $23k per year (Estimated) • In office • Internship • PhD • Shanghai
C++
Python
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenCL
DevOps
HPC
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
C#
C++
Python
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
AI Agents
Apply
$13k – $21k per year (Estimated) • In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$83k – $124k per year • In office • 3+ years exp • Bachelor's Degree • Austin
Design
AutoCAD
Management
Microsoft Project
Apply
$83k – $124k per year • In office • 3+ years exp • Bachelor's Degree • Austin
Design
AutoCAD
Management
Microsoft Project
Apply
$108k – $210k per year (Estimated) • In office • 2+ years exp • PhD • Austin
Design
AutoCAD
Apply
$131k – $299k per year (Estimated) • In office • 7+ years exp • Master's Degree • Austin
Apply
$56k – $94k per year • In office • Contractor • Austin
Apply
See all jobs
This is one of many
413,737 more open roles from verified company boards, updated every day.