406,377open jobs
14,133companies
78,486added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 10+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are looking for a Senior Software Engineer to help build NeMo Platform, NVIDIA’s product for developing, evaluating, deploying, and operating AI systems at scale. This role is for a senior engineer/architect for our Core team which owns and ships an open source plugin-based AI platform for running and optimizing Agents targeting multiple compute backends (local/docker, Kubernetes, Slurm, etc.).

As AI systems become more autonomous and more deeply integrated into real workflows, teams need robust APIs and orchestration systems for running, monitoring, and optimizing Agents at scale. Increasingly, the users of these systems are themselves autonomous or semi-autonomous Agents. The NeMo Platform group is building a sophisticated agent execution framework to enable agents to automatically run hundreds of experiments in parallel to find the most efficient agent architectures for our customers. This is important product engineering research for making agents more sustainable. AI systems are not yet nearly as efficient as they can be, and systems like NeMo Platform will allow large scale AI consumers to automatically tune their agents to use fewer tokens and rely on more efficient models with better throughput.

What you'll be doing:

  • Working in a product research environment where we place big bets on where the future is heading, adapting in real time as we build alongside an industry that is constantly evolving with us. This means fast iteration, high ownership, pragmatic decisions, and performance-minded implementation under production constraints

  • Designing an Agentic Execution system that are flexible enough to work in many environments (local, k8s, Slurm, on prem / air-gapped)

  • Provide senior technical leadership through design reviews, code reviews, mentoring, and ownership of ambiguous cross-component problems

  • Building and maintain our Core Platform APIs for running jobs, storing data, entities, secrets, and RBAC and Auth

  • Extending our flexible Plugin Architecture that makes it easy for many teams and external customers to install new capabilities into our system

  • Building in the open in our OSS repo, keeping up the high standards that the open source community demands

  • Shipping code at the speed of light with an unlimited token budget using best in class agentic coding tools

  • Improving reliability, observability, debuggability, and performance across NeMo Platform, SDKs, plugins, jobs, and developer workflows

  • Building strong test coverage across unit, integration, E2E, Docker, and Kubernetes workflows

What we need to see:

  • BS, MS, or equivalent experience in Computer Science, Computer Engineering, or a related technical field

  • 10+ years of professional software engineering experience building production systems

  • Comfort working in a very fast and ambiguous environment

  • Exceptional communication, both verbal and written. This includes the ability to produce and review high quality architectural RFCs, and to discuss them clearly with the right level of technical detail for the right people (Engineer, Product, Marketing, etc.)

  • Strong system design skills, with a pragmatic flexibility and phenomenal instincts to invent robust systems quickly without over-complicating. Strong understanding of reliability, scalability, security, and performance tradeoffs in production infrastructure

  • Experience with distributed systems, cloud-native services, containers, Kubernetes, and job orchestration

  • Excellent Python engineering skills, including API design, typing, testing, debugging, performance analysis, and maintainable software design

  • Experience designing SDKs, libraries, plugins, CLIs, or other developer-facing interfaces

  • Ability to work independently, define technical scope, break down ambiguous problems, and drive work across team boundaries

Ways to stand out from the crowd:

  • Experience building, deploying, and iterating on production agentic AI systems at scale in Kubernetes

  • Experience with sophisticated plugin architectures

  • Strong ability to connect technical evaluation work to business outcomes, product quality, user experience, reliability, or operational efficiency

  • Experience with enterprise AI systems where measurement, regression testing, observability, governance, and continuous improvement are required for production deployment

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 4, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
406,377 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
AI Lead 4 days ago
$35k – $83k per year (Estimated) • In office • Bengaluru
Python
Python
FastAPI
Flask
Databases
Chroma
Databricks
FAISS
Pinecone
Weaviate
AI/ML
AI Agents
Amazon SageMaker
AutoGen
AWS Bedrock
Claude
CrewAI
Embeddings
Fine-tuning
Function Calling
Gemini
Hugging Face
Kubeflow
LangChain
LangGraph
LlamaIndex
LLM
LLMOps
LoRA
MLFlow
NLP
OpenAI
PEFT
Prompt Engineering
PyTorch
RAG
Semantic Kernel
Semantic Search
Semantic Search
TensorFlow
Transformers
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Rest API
Vector
Analytics
ETL/ELT
Apply
$129k – $178k per year • Remote • Full-Time
Python
TypeScript
JavaScript
Databases
ElasticSearch
PostgreSQL
Frontend
Ant Design
React.js
DevOps
AWS
Datadog
Docker
Kubernetes
Terraform
Apply
$8.5k – $26k per year (Estimated) • Remote • Full-Time • 2+ years exp • Moscow
PHP
Python
Databases
PostgreSQL
DevOps
Docker
Git
Rest API
Apply
$107k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
SQL
Databases
MySQL
DevOps
Ansible
AppDynamics
AWS
Blue-Green Deployment
CI/CD
Datadog
Docker
Dynatrace
Grafana
Jenkins
Kubernetes
New Relic
Prometheus
Terraform
Robotics
Digital Twin
Apply
$186k – $227k per year • Equity • Remote • Full-Time • 7+ years exp
Databases
PostgreSQL
Snowflake
DevOps
Azure
CI/CD
Datadog
GitOps
Grafana
Kubernetes
Platform Engineering
Cybersecurity
SOC 2
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
$107k – $269k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tel Aviv
SQL
Databases
MySQL
DevOps
Ansible
AppDynamics
AWS
Blue-Green Deployment
CI/CD
Datadog
Docker
Dynatrace
Grafana
Jenkins
Kubernetes
New Relic
Prometheus
Terraform
Robotics
Digital Twin
Apply
$63k – $151k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tokyo
C++
Python
C++
PyTorch C++
AI/ML
CUDA
CUDA Toolkit
PyTorch
DevOps
HPC
Apply
$221k – $387k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
DevOps
Incident Management
Management
ServiceNow
Apply
$56k – $94k per year • In office • Contractor • Santa Clara
Apply
$64k – $74k per year • In office • Contractor • Santa Clara
Apply
$56k per year • In office • Contractor • Santa Clara
Apply
$84k – $94k per year • In office • Contractor • Santa Clara
Apply
See all jobs
This is one of many
406,377 more open roles from verified company boards, updated every day.