368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$152k – $242k per year
Location
In office (Santa Clara, Seattle)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPUs act as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent.

As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. NVIDIA is widely recognized as one of the most desirable employers, with some of the most versatile people in the world working for us. If you're passionate about building scalable, efficient backend systems to power cloud operations, we invite you to join our team.

We are looking for a Senior Software Engineer to join our DGX Cloud / Fleet Intelligence team and build backend systems that power GPU health monitoring, telemetry ingestion, operational automation, and fleet visibility for NVIDIA’s high-performance GPU infrastructure.

What You'll Be Doing:

This role focuses on backend services for GPU health, Fleet Intelligence, telemetry ingestion, inventory, attestation, alerting, reporting, and cloud operations automation, in addition, you will:

  • Design and develop Go backend services, REST APIs, and data models for GPU Health and Fleet Intelligence.

  • Build customer-facing and agent-facing APIs for compute zones, node groups, nodes, alerts, events, metrics, inventory, attestation, retention policies, and reports.

  • Develop high-volume ingestion and persistence paths for in-band agents and out-of-band collectors.

  • Work with Aurora PostgreSQL, Kafka/MSK, S3, SQS, Prometheus/AMP, and OpenTelemetry-based observability.

  • Build and operate scheduled backend services for liveness tracking, alerting, rollups, cleanup, notification delivery, attestation, and XID analysis.

  • Optimize database schemas, partitioned time-series storage, query performance, CTE-heavy queries, and connection pooling for reliable service behavior.

  • Collaborate with agent, infrastructure, SRE, UI, and cloud operations teams to turn operational workflows into scalable backend systems.

  • Improve service reliability, security, observability, testing, and deployment quality across Docker, Kubernetes, Helm, and cloud environments.

What We Need To See:

  • 5+ years of industry software engineering experience with a Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or equivalent experience.

  • Strong Go backend development experience.

  • Experience building REST APIs and production services using frameworks such as Gin or similar.

  • Strong PostgreSQL experience, including schema design, query tuning, transactions, migrations, and operational data modeling.

  • Experience with distributed systems, event pipelines, telemetry ingestion, or operational analytics.

  • Experience with cloud infrastructure, Docker, Kubernetes, and Helm.

  • Strong debugging skills across APIs, databases, cloud services, and production systems.

  • Familiarity with authentication and authorization patterns such as JWT, service account keys, trusted-edge proxies, or customer-scoped APIs.

  • Experience with Linux-based development and production environments.

Ways To Stand Out From The Crowd:

  • Background with telemetry, monitoring, observability, health, or fleet-management platforms.

  • Experience with AWS services such as MSK, Aurora, S3, SQS, AMP, or CloudWatch.

  • Experience with OpenTelemetry, Prometheus, LightStep, Grafana, or production tracing/metrics systems.

  • Background with NVIDIA GPUs, DGX systems, DCGM, XID analysis, attestation, or AI datacenter operations.

  • Experience designing APIs and storage systems that support real-time operational workflows at fleet scale.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 28, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$217k – $304k per year • Equity • Remote • Full-Time • 8+ years exp
Go
Databases
Apache Kafka
ClickHouse
Google BigQuery
AI/ML
Flink
Recommender Systems
DevOps
Incident Management
Kubernetes
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$185k – $260k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
DevOps
AWS
CI/CD
GCP
Kubernetes
GitHub
Cybersecurity
Clair
Dependabot
OWASP Top 10
OWASP ZAP
Snyk
Trivy
Apply
$110k – $131k per year • Remote • Full-Time • 10+ years exp • Bachelor's Degree
DevOps
AWS
Incident Management
VMWare
Apply
$169k – $321k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Phoenix
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
AWS
Kong
Amazon S3
API Gateway
Cybersecurity
Zero Trust
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$156k – $348k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • United Kingdom
AI/ML
CUDA
CUDA Toolkit
AI Agents
NVIDIA NeMo
Apply
$72k – $99k per year • Equity • In office • Full-Time • Santa Clara
Apply
$166k – $290k per year • Equity • In office • Full-Time • 8+ years exp • Santa Clara
Management
ServiceNow
Apply
$133k – $272k per year (Estimated) • In office • Santa Clara
Go
Python
AI/ML
Edge AI
LLM
RAG
Apply
$80k – $110k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Apply
$142k – $256k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara • Toronto
C++
Go
IoT
MQTT
OPC UA
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.