514,690open jobs
17,495companies
76,717added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA seeks a Senior Software Engineer specializing in Deep Learning Inference for our growing team. As a key contributor, you will help design, build, and optimize the GPU-accelerated software that powers today’s most sophisticated AI applications. Our team is responsible for developing and maintaining high-performance deep learning frameworks, including SGLang and vLLM, which are at the forefront of efficient large-scale model serving and inference. You will play a central role in improving these platforms, facilitating smooth deployment and serving of groundbreaking language models.

You’ll work closely with the deep learning community to implement the latest algorithms for public release in frameworks like SGLang and vLLM, as well as other DL frameworks. Your work will focus on identifying and driving performance improvements for state-of-the-art LLM and Generative AI models across NVIDIA accelerators, from datacenter GPUs to edge SoCs. You'll bring to bear open-source tools and plugins-including CUTLASS, OAI Triton, NCCL, and CUDA kernels-to implement and optimize model serving pipelines.

What you'll be doing:

  • Performance optimization, analysis, and tuning of DL models in various domains like LLM, Multimodal and Generative AI.

  • Scale performance of DL models across different architectures and types of NVIDIA accelerators.

  • Contribute features and code to NVIDIA’s inference libraries, vLLM and SGLang, FlashInfer and LLM software solutions.

  • Work with cross-collaborative teams across frameworks, NVIDIA libraries and inference optimization innovative solutions.

What we need to see:

  • Masters or PhD or equivalent experience in relevant field (Computer Engineering, Computer Science, EECS, AI).

  • 6+ years of relevant software development experience.

  • Excellent C/C++ programming and software design skills. SW Agile skills are helpful and Python experience is a plus.

  • Prior experience with training, deploying or optimizing the inference of DL models in production is a plus.

  • Prior background with performance modeling, profiling, debug, and code optimization or architectural knowledge of CPU and GPU is a plus.

Ways to stand out from the crowd:

  • Contribute to Deep Learning Software projects, such as PyTorch, vLLM, and SGLang to drive advancements in the field.

  • Experience with Multi-GPU Communications (NCCL, NVSHMEM)

  • Experience building and shipping products to enterprise customers.

  • GPU programming experience (CUDA, OAI TRITON or CUTLASS).

NVIDIA is at the forefront of breakthroughs in Artificial Intelligence, High-Performance Computing, and Visualization. Our teams are composed of driven, innovative professionals dedicated to pushing the boundaries of technology.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 14, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
514,690 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$24k – $58k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
Python
AI/ML
CUDA Toolkit
Multimodal AI
Function Calling
Computer Vision
AI Agents
TensorRT
Semantic Search
CUDA
Semantic Search
Tool Use
Physical AI
DevOps
Helm
GitHub Actions
GitLab CI
CI/CD
Jenkins
Docker
Kubernetes
Apply
Engineering Intern 1 hour ago
$38k – $58k per year (Estimated) • In office • Internship • Bachelor's Degree • Princeton
Python
C#
C++
DevOps
Git
Apply
$30k – $79k per year (Estimated) • In office • Full-Time • 4+ years exp • Master's Degree • Shanghai
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorRT
PyTorch
LLM
CUDA
cuDNN
Apply
$75k – $114k per year • In office • Top Secret • Full-Time • 3+ years exp • Bachelor's Degree • Albuquerque • Dayton • Simi Valley • Herndon • Arlington
Python
PowerShell
Cybersecurity
Microsoft Sentinel
VirusTotal
MITRE ATT&CK
Cyber Kill Chain
KEV
Tanium
Microsoft Entra ID
Apply
$91k – $139k per year • In office • Top Secret • Full-Time • 3+ years exp • Bachelor's Degree • Albuquerque • Dayton • Simi Valley • Herndon • Arlington
Python
SQL
PowerShell
DevOps
Rest API
Cybersecurity
Microsoft Defender
CIS Benchmarks
CVSS
KEV
Tanium
Apply
$24k – $58k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune
Python
AI/ML
CUDA Toolkit
Multimodal AI
Function Calling
Computer Vision
AI Agents
TensorRT
Semantic Search
CUDA
Semantic Search
Tool Use
Physical AI
DevOps
Helm
GitHub Actions
GitLab CI
CI/CD
Jenkins
Docker
Kubernetes
Apply
$192k – $305k per year • In office • Full-Time • Master's Degree • Redmond • Santa Clara
AI/ML
CUDA Toolkit
CUDA
Apply
$124k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • PhD • Santa Clara
Apply
$136k – $219k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
DevOps
HPC
Chips/EDA
Cadence Allegro
Apply
$30k – $79k per year (Estimated) • In office • Full-Time • 4+ years exp • Master's Degree • Shanghai
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorRT
PyTorch
LLM
CUDA
cuDNN
Apply
$46k – $71k per year (Estimated) • In office • Internship • PhD • Santa Clara • Irvine • Austin • Morrisville • Chandler
Python
AI/ML
Copilot
LangGraph
AutoGen
LangChain
Claude
ChatGPT
LlamaIndex
Model Context Protocol
Fine-tuning
Reinforcement Learning
Prompt Engineering
Multimodal AI
Diffusion Models
AI Agents
TensorFlow
PyTorch
CrewAI
LLM
RAG
Hugging Face
A2A
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Git
Management
n8n
Apply
$61k – $100k per year • In office • Bachelor's Degree • Santa Clara
Apply
$167k – $291k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
Bash
Databases
Azure Cosmos DB
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
Terraform
GCP
CloudFormation
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Self-Healing
Amazon EKS
Google GKE
Azure AKS
AWS Lambda
Amazon EC2
AIOps
Incident Management
SLI/SLO/SLA
Management
ServiceNow
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Apply
$168k – $270k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
AI/ML
Spark
DevOps
Terraform
Azure
AWS
Docker
Kubernetes
Analytics
Tableau
Power BI
ETL/ELT
SAP BusinessObjects
Management
Outlook
Apply
See all jobs
This is one of many
514,690 more open roles from verified company boards, updated every day.