401,286open jobs
13,975companies
77,769added this week
Browse all
Salary
$184k – $288k per year
Location
In office (United States, Seattle, Austin, Santa Clara, Redmond)
Seniority
Senior · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are developing advanced multi-rack, multi-tenant AI/ML datacenters with NVIDIA GB200, and upcoming GB300 GPUs. NVIDIA seeks a Senior Software Engineer for our CSP (Cloud Service Provider) Engagements team to focus on the cloud-native stack for datacenter products like GB200. In this role, You will define customer workflows, prototype stack enhancements, and debug the toughest Kubernetes + Slurm issues in multi-rack, multi-tenant AI datacenters. You'll tackle complex scheduling challenges across racks, tenants, and clouds as part of the CSP engagements team.

What you’ll be doing:

  • Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies, etc.

  • Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities.

  • Drive joint architecture reviews and “whiteboard” sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests.

  • Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites.

  • Deliver technical collateral-design docs, how-to guides, demo scripts-and present at customer on-sites, KubeCon, and SlurmUG.

  • Collaborate with AE, FAE, and Solution Architect teams to deliver integrated customer solutions and technical documentation.

What we need to see:

  • Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins).

  • Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters.

  • Proven track record debugging large-scale, cloud-native stacks across networking (RDMA/RoCE), storage, and control planes.

  • Customer-facing engineering or solutions-architect background: requirements gathering, PoC ownership, roadmap influence.

  • Familiarity with CI/CD (GitHub Actions, Tekton), observability (Prometheus, OpenTelemetry), and infrastructure-as-code.

  • Excellent communication-able to switch between deep technical detail and high-level business impact.

  • 10+ years of professional software development experience in distributed systems (Go, Rust, C/C++ or Python for tooling).

  • BS or MS (or equivalent experience) in Computer Engineering, Computer Science, or related field.

Ways to stand out from the crowd:

  • Upstream contributions to Kubernetes, Slurm, Volcano, or similar projects.

  • Experience with GPU computing (CUDA), deep learning workloads

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. We have some of the most forward-thinking and hardworking people in the world working for us. If you're creative, hardworking and self-motivated, we want to hear from you!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 5, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
401,286 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seattle
$14k – $35k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Phoenix
Bash
PowerShell
Python
AI/ML
AI Agents
Edge AI
DevOps
Ansible
Azure
Azure AKS
Azure DevOps
Bamboo
Blue-Green Deployment
Chef
CI/CD
Configuration Management
Datadog
Docker
Git
GitHub
GitHub Actions
GitLab
Grafana
Incident Management
Jenkins
Kubernetes
Platform Engineering
Splunk
Terraform
Cybersecurity
Microsoft Entra ID
Cryptography
Vault
Marketing
LinkedIn
YouTube
Apply
$60k – $300k per year • Equity 0.1–0.5% • Remote • Full-Time • 6+ years exp • San Francisco
C++
Elixir
JavaScript
Kotlin
Node JS
Python
Rust
Swift
TypeScript
AI/ML
CUDA
CUDA Toolkit
Edge AI
Embeddings
Semantic Search
Semantic Search
Frontend
npm
WebAssembly
DevOps
Vector
Apply
Sr. DevOps Engineer 2 days ago
$108k – $210k per year (Estimated) • Remote • 5+ years exp • PhD • Reston
Databases
Apache Kafka
Kafka
Oracle
PostgreSQL
DevOps
Amazon ECS
Amazon EKS
Ansible
AWS
CI/CD
CloudFormation
Git
GitHub
Kubernetes
New Relic
Splunk
Terraform
Ubuntu
Management
Confluence
Microsoft Teams
Apply
$91k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Ann Arbor
Python
DevOps
Ansible
Azure
CI/CD
Configuration Management
Terraform
Cybersecurity
Palo Alto NGFW
Apply
$160k – $250k per year • In office • Full-Time • Bachelor's Degree • Dallas
C++
DevOps
CI/CD
RTOS
Apply
$272k – $431k per year • In office • Full-Time • 15+ years exp • PhD • Santa Clara • New York
AI/ML
AI Agents
Fine-tuning
Function Calling
LLM
Multimodal AI
NVIDIA NeMo
Post-training
Pre-training
Reinforcement Learning
Structured Outputs
Synthetic Data
TGI
vLLM
DevOps
CI/CD
Git
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
$144k – $230k per year • In office • Full-Time • PhD • Santa Clara
Apex
Apex
MuleSoft
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
In office • Full-Time • 5+ years exp • Master's Degree • Shanghai
C++
Python
AI/ML
AI Agents
Copilot
Cursor
LLM
Reinforcement Learning
Robotics
Isaac Lab
Isaac Sim
Perception
Reinforcement Learning
ROS
Sim-to-Real
Teleoperation
Apply
$250k – $300k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Seattle
DevOps
AWS
GCP
Marketing
Salesforce
Apply
$120k – $250k per year • Equity 0.2–1% • In office • Full-Time • Master's Degree • Seattle
Python
AI/ML
Diffusion Models
Fine-tuning
PyTorch
Self-Supervised Learning
Apply
$266k – $445k per year • Remote/Hybrid • Full-Time • San Francisco • Seattle
JavaScript
AI/ML
OpenAI
Frontend
Bootstrap
DevOps
Platform Engineering
Apply
Senior Engineer 1 day ago
$200k – $260k per year • Equity 0.2–0.2% • In office • Full-Time • 6+ years exp • Seattle
Python
TypeScript
Python
Django
AI/ML
AI Agents
LLM
Edge AI
Model Context Protocol
Mobile
React Native
Apply
Product Manager 1 day ago
$180k – $250k per year • Equity 0.2–0.2% • In office • Full-Time • 6+ years exp • Seattle
SQL
AI/ML
Edge AI
Apply
See all jobs
This is one of many
401,286 more open roles from verified company boards, updated every day.