368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$124k – $196k per year
Location
In office (Santa Clara, United States)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is recruiting a Senior Inference Performance Engineer to push NVIDIA's performance limits on large-scale AI inference benchmarks. This position provides an outstanding opportunity to employ your optimization knowledge in an autonomous optimization framework. AI agents use this framework to repeatedly run benchmark, profile, and tune processes, amplifying the impact of every technique you design. If you enjoy extracting maximum performance from GPUs and scaling your skills beyond your individual efforts, this role is a great fit!

What you'll be doing:

  • Distill your performance instincts into reusable skills, workflows, and evidence-backed methodologies that AI agents can complete autonomously. Review agent-generated experiments, validate findings, and curate best-known configurations.

  • Performance improvement of AI inference workloads that methodically increase throughput-per-GPU and user interactivity by exploring configuration options, parallelism techniques, batching, KV cache handling, quantization, and speculative decoding settings.

  • Measure and optimize both aggregated and disaggregated serving architectures across TensorRT-LLM, SGLang, vLLM, and Dynamo on NVIDIA's latest GPU platforms.

  • Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.

  • Land improvements upstream: serving framework patches, optimized kernels, and deployment recipes that advance the public Pareto frontier while maintaining strict model correctness.

  • Collaborate with TensorRT-LLM, SGLang, vLLM, kernel, benchmarking, and GPU architecture teams to convert profiling insights into delivered performance improvements.

What we need to see:

  • BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Applied Math, or a related field, or equivalent experience.

  • 3+ years of relevant engineering experience.

  • Must have: Extensive knowledge of the efficiency and optimization involved in AI model execution, covering continuous batching, throughput-latency tradeoffs, KV cache and memory limitations, parallel processing techniques, MoE serving, quantization, and meeting serving SLAs.

  • Must have: Hands-on experience benchmarking and profiling GPU workloads using tools such as Nsight Systems, Nsight Compute, CUPTI, or PyTorch profiler, and interpreting kernel-level performance data.

  • Strong Python engineering skills and the ability to navigate and modify large C++/CUDA serving codebases.

  • Rigorous experimental methodology with controlled single-variable comparisons, reproducible benchmarks, and evidence-backed optimization decisions.

  • Strong written and verbal communication skills to explain performance tradeoffs clearly to both humans and documentation for autonomous systems.

Ways to stand out from the crowd:

  • Direct contributions to TensorRT-LLM, vLLM, SGLang, FlashInfer, Dynamo, or comparable inference frameworks.

  • Experience with disaggregated serving, wide expert-parallel MoE inference, KV cache transfer, or NCCL/NIXL/NVSHMEM communication at multi-node scale.

  • CUDA kernel authorship or optimization experience on Hopper/Blackwell architectures, focusing on Tensor Cores, TMA, and warp specialization.

  • Proven results on public inference benchmarks such as MLPerf Inference or SemiAnalysis InferenceX.

  • Experience building or operating agentic AI workflows to automate engineering tasks.

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 124,000 USD - 195,500 USD for Level 2, and 152,000 USD - 241,500 USD for Level 3.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 11, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$26k – $66k per year (Estimated) • In office • Full-Time • 7+ years exp • Master's Degree • Hyderabad
Python
SQL
TypeScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
AWS Bedrock
Claude
Claude Code
Fine-tuning
Hallucination
LangChain
LLM
Model Context Protocol
Multimodal AI
Prompt Engineering
RAG
Synthetic Data
A2A
Amazon SageMaker
DevOps
AWS
CI/CD
Docker
Vector
Analytics
A/B Testing
Apply
$22k – $57k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hyderabad
Python
TypeScript
JavaScript
Databases
OpenSearch
Snowflake
AI/ML
AI Agents
Claude
Claude Code
LLM
Model Context Protocol
Frontend
Angular
Next.js
React.js
DevOps
AWS
AWS Lambda
CI/CD
Docker
GitHub Actions
Terraform
Amazon ECS
Amazon S3
GitHub
IAM
Apply
AI Engineer 2 days ago
Remote/Hybrid • 8+ years exp
C#
Python
TypeScript
C#
.NET
Databases
ElasticSearch
OpenSearch
AI/ML
AutoGen
AWS Bedrock
Claude
CrewAI
Embeddings
LangGraph
LLM
Model Context Protocol
Prompt Engineering
RAG
Semantic Kernel
Semantic Search
LangChain
Anthropic
LLM Guardrails
OpenAI
OpenAI Agents SDK
Semantic Search
Structured Outputs
AI Agents
Function Calling
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Vector
Apply
$140k – $253k per year (Estimated) • In office • 5+ years exp • Long Beach
C++
Rust
C++
Protobuf
AI/ML
Human-in-the-Loop
DevOps
CI/CD
Vector
Apply
In office • Full-Time • 3+ years exp • Master's Degree • Shanghai • Shenzhen
C++
C
C++
TBB
C
MPI
Pthreads
AI/ML
CUDA
CUDA Toolkit
cuDF
OpenMP
RAPIDS
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Perl
Python
Apply
$98k – $252k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Switzerland
Assembly
C++
Fortran
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenMP
DevOps
HPC
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Hsinchu
Perl
Python
Apply
In office • Full-Time • 5+ years exp • Hsinchu • Taipei
C++
Python
AI/ML
InfiniBand
Apply
$100k – $137k per year • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Apply
$72k – $99k per year • Equity • In office • Full-Time • Santa Clara
Apply
$166k – $290k per year • Equity • In office • Full-Time • 8+ years exp • Santa Clara
Management
ServiceNow
Apply
$133k – $272k per year (Estimated) • In office • Santa Clara
Go
Python
AI/ML
Edge AI
LLM
RAG
Apply
$80k – $110k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.