1,432,387open jobs
84,237companies
218,256added this week
Browse all
Salary
$184k – $288k per year
Location
In office (Santa Clara, United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 9, 2026. NVIDIA scores A on the Alion truth index.

Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is at the forefront of the generative AI revolution! The Algorithmic Model Optimization Team specifically focuses on optimizing generative AI models such as large language models (LLM) and diffusion models for maximal inference efficiency using techniques ranging from neural architecture search and pruning to sparsity, quantization, and automated deployment strategies. Our work includes conducting applied research to improve model efficiency as well as developing an innovative software platform (TRT Model Optimizer). Our software is used both internally across NVIDIA and externally by research and engineering teams alike developing best-in-class AI models.

We are now looking for a Senior Deep Learning Software Engineer to develop and scale up our automated inference and deployment solution. As part of the team, you will be instrumental in pushing the limits of inference efficiency and large-scale, automated deployment. Your work will touch upon fundamental aspects of a typical machine learning stack including working in high-level frameworks like PyTorch and HuggingFace to developing and improving high-performance kernel implementations in CUDA, TRT-LLM, and Triton. This is an exceptional opportunity for passionate software engineers straddling the boundaries of research and engineering, with a strong background in both machine learning fundamentals and software architecture & engineering.

What you’ll be doing:

  • Train, develop, and deploy state-of-the generative AI models like LLMs and diffusion models using NVIDIA's AI software stack.
  • Leverage and build upon the torch 2.0 ecosystem (TorchDynamo, torch.export, torch.compile, etc...) to analyze and extract standardized model graph representation from arbitrary torch models for our automated deployment solution.
  • Develop high-performance optimization techniques for inference, such as automated model sharding techniques (e.g. tensor parallelism, sequence parallelism), efficient attention kernels with kv-caching, and more.
  • Collaborate with teams across NVIDIA to use performant kernel implementations within our automated deployment solution.
  • Analyze and profile GPU kernel-level performance to identify hardware and software optimization opportunities.
  • Continuously innovate on the inference performance to ensure NVIDIA's inference software solutions (TRT, TRT-LLM, TRT Model Optimizer) can maintain and increase its leadership in the market.
  • Play a pivotal role in architecting and designing a modular and scalable software platform to provide an excellent user experience with broad model support and optimization techniques to increase adoption.

What we need to see:

  • Masters, PhD, or equivalent experience in Computer Science, AI, Applied Math, or related field.
  • 8+ years of relevant work or research experience in Deep Learning.
  • Excellent software design skills, including debugging, performance analysis, and test design.
  • Strong proficiency in Python, PyTorch, and related ML tools (e.g. HuggingFace).
  • Strong algorithms and programming fundamentals.
  • Good written and verbal communication skills and the ability to work independently and collaboratively in a fast-paced environment.

Ways to stand out from the crowd:

  • Contributions to PyTorch, JAX, or other Machine Learning Frameworks.
  • Knowledge of GPU architecture and compilation stack, and capability of understanding and debugging end-to-end performance.
  • Familiarity with NVIDIA's deep learning SDKs such as TensorRT.
  • Prior experience in writing high-performance GPU kernels for machine learning workloads in frameworks such as CUDA, CUTLASS, or Triton.

Increasingly known as “the AI computing company” and widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. Are you creative, motivated, and love a challenge? If so, we want to hear from you! Come, join our model optimization group, where you can help build real-time, cost-effective computing platforms driving our success in this exciting and rapidly-growing field.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 184,000 USD - 287,500 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until October 13, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,432,387 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Santa Clara
$180k – $250k per year • In office • Full-Time • 7+ years exp • Palo Alto
JavaScript
TypeScript
AI/ML
World Models
Frontend
React.js
Apply
$165k – $200k per year • In office • 7+ years exp • Chicago
Python
AI/ML
LangGraph
AutoGen
LangChain
LlamaIndex
AI Agents
CrewAI
LLM
DevOps
GCP
Azure
CI/CD
AWS
Management
Agile
Apply
$104k – $130k per year • Hybrid • 5+ years exp • Chicago
Python
SQL
AI/ML
Hadoop
Spark
Reinforcement Learning
Scikit-learn
Computer Vision
NLP
TensorFlow
PyTorch
Interpretability
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
A/B Testing
Management
Agile
Apply
$117k – $227k per year • In office • Full-Time • 12+ years exp • PhD • El Segundo
Python
MATLAB
AI/ML
Machine Learning
Apply
$104k – $218k per year • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Ashburn
Python
AI/ML
OpenCV
CUDA Toolkit
YOLO
Fine-tuning
Computer Vision
ONNX
TensorRT
Label Studio
TensorFlow
PyTorch
Ultralytics
CUDA
Amazon SageMaker
CVAT
Roboflow
cuDNN
DevOps
AWS
Amazon EC2
Management
Agile
Scrum
Apply
$135k per year • Remote (United States) • Full-Time • 5+ years exp
Python
C++
DevOps
RTOS
CI/CD
AWS
Wi-Fi
IoT
MQTT
Zigbee
ESP-IDF
Apply
≈ $101k – $206k per year (Estimated) • Remote (Canada) • Full-Time • 5+ years exp
Python
C++
DevOps
RTOS
CI/CD
AWS
Wi-Fi
IoT
MQTT
Zigbee
ESP-IDF
Apply
≈ $118k – $255k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp
Python
C++
DevOps
RTOS
CI/CD
AWS
Wi-Fi
IoT
MQTT
Zigbee
ESP-IDF
Apply
≈ $118k – $255k per year (Estimated) • Remote (Canada) • Full-Time • 8+ years exp
Python
C++
DevOps
RTOS
CI/CD
AWS
Wi-Fi
IoT
MQTT
Zigbee
ESP-IDF
Apply
≈ $117k – $239k per year (Estimated) • Equity • Remote (United States) • 10+ years exp
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
AI/ML
Copilot
Cursor
Claude
dbt
Flink
Anomaly Detection
Feature Store
Edge AI
Machine Learning
DevOps
Terraform
CI/CD
AWS
SLI/SLO/SLA
Cybersecurity
GDPR
Analytics
Dimensional Modeling
Apply
$224k – $357k per year • In office • Full-Time • 12+ years exp • PhD • Santa Clara
Python
Go
Java
AI/ML
Cursor
Claude
Model Context Protocol
Fine-tuning
Function Calling
AI Agents
RAG
OpenAI Codex
Tool Use
Machine Learning
DevOps
Rest API
Terraform
Ansible
GitLab CI
CI/CD
Jenkins
Apply
$152k – $242k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara • Austin • Redmond
Python
C++
C++
PyTorch C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
OpenCL
PyTorch
CUDA
Recommender Systems
MLIR
Apache TVM
XLA
Apply
$76k – $188k per year • In office • Internship • PhD • Santa Clara
Python
AI/ML
Function Calling
AI Agents
PyTorch
Tool Use
Apply
$136k – $219k per year • In office • Full-Time • 5+ years exp • PhD • Durham
Python
Verilog
C++
SystemVerilog
Chips/EDA
UVM
Apply
$120k – $207k per year • In office • Full-Time • PhD • Santa Clara • Seattle
Python
Bash
AI/ML
Cursor
Claude
AI Agents
OpenAI Codex
NCCL
InfiniBand
NVLink
DevOps
SLURM
AWS
KVM
HPC
Linux
Cybersecurity
Wireshark
Tcpdump
Apply
≈ $223k – $492k per year (Estimated) • In office • PhD • Santa Clara
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Fine-tuning
Multimodal AI
Diffusion Models
Computer Vision
TensorFlow
PyTorch
GAN
Machine Learning
Apply
≈ $217k – $413k per year (Estimated) • In office • Santa Clara
Apply
≈ $215k – $476k per year (Estimated) • In office • Santa Clara
AI/ML
Embeddings
Machine Learning
Apply
≈ $263k – $537k per year (Estimated) • In office • Santa Clara
AI/ML
Machine Learning
Apply
≈ $227k – $502k per year (Estimated) • In office • Santa Clara
AI/ML
Multimodal AI
LLM
Apply
See all jobs
This is one of many
1,432,387 more open roles from verified company boards, updated every day.