407,354open jobs
14,154companies
73,129added this week
Browse all
Salary
$192k – $305k per year
Location
Remote (United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

We are now looking for a Senior Research Engineer passionate about Generative AI inference. Are you excited to change the way people infuse AI into products and services? NVIDIA is at the forefront of generative AI models, from language to images. NVIDIA provides building blocks to democratize AI and make generative AI easy to develop, integrate, and deploy. Our team is dedicated to developing optimized inferencing technologies to support our growing generative AI needs. We contribute to all steps of the machine learning lifecycle: from conceptualization, to applied research, engineering for optimized inference, and deployment. Collaborate with research teams, engineers, and open-source community.

What you will be doing:

  • Design and evaluate routing policies for LLM traffic to best use mixture of model systems.

  • Build and run agentic benchmarks (e.g., Terminal-Bench ) to measure algorithm quality, and turn results into calibration data and routing profiles

  • Ship to an open-source repo: design docs, code review, docs, and community contributions

  • Collaborating with engineering teams across all of NVIDIA to ensure our software integrates seamlessly up and down the NVIDIA accelerated serving stack.

What we need to see:

  • Bachelor's of Master's degree in Computer Science or equivalent experience.

  • 8+ years of industry experience in Deep Learning frameworks (PyTorch or TensorFlow).

  • Experience designing or running LLM evaluations/benchmarks - ideally agentic ones - and drawing statistically sound conclusions from them

  • Understanding of modern techniques in Machine Learning, Deep Neural Networks, Natural Language Processing, or Speech Recognition.

  • Empirical research mindset: forming hypotheses about new algorithms, running calibrations, iterating on results

  • Strong communication and interpersonal skills, along with the ability to work in a dynamic and distributed team. A history of mentoring junior engineers and interns is a huge plus.

  • A desire to constantly grow and learn new things.

  • Strong computer science fundamentals - algorithms and data structures, computational complexity, parallel and distributed computing, system software.

Ways to stand out from a crowd:

  • Experience architecting or developing large-scale distributed systems for deep learning.

  • Agentic benchmark creation and publications.

  • Knowledge of CPU and/or GPU architecture.

  • GPU programming (CUDA).

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 192,000 USD - 304,750 USD for Level 4, and 224,000 USD - 356,500 USD for Level 5.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until July 14, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
407,354 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
ML engineer 2 days ago
$81k – $116k per year • Equity 0.1–0.5% • Remote/Hybrid • Full-Time • 6+ years exp • Paris
Python
Rust
TypeScript
AI/ML
AI Agents
Claude
JAX
Multimodal AI
PyTorch
Reinforcement Learning
SGLang
TensorFlow
vLLM
DevOps
CloudFormation
GCP
Terraform
Robotics
Digital Twin
Apply
Product Engineer 1 day ago
$150k – $225k per year • Equity 0.2–0.4% • Remote • Full-Time • 3+ years exp • San Francisco
Python
Python
FastAPI
AI/ML
AI Agents
LLM
Frontend
Next.js
DevOps
Azure
WebSockets
Apply
Product Manager 4 days ago
$19k – $50k per year (Estimated) • In office • 2+ years exp • Noida
AI/ML
AI Agents
Apply
Lead data engineer 2 days ago
$93k – $116k per year • Equity 0.3–1% • Remote/Hybrid • Full-Time • 11+ years exp • Paris
Python
Rust
SQL
TypeScript
Databases
Apache Kafka
BigQuery
Google BigQuery
Kafka
Snowflake
AI/ML
AI Agents
Claude
Dagster
dbt
Multimodal AI
Reinforcement Learning
DevOps
CloudFormation
GCP
Terraform
Robotics
Digital Twin
Analytics
ETL/ELT
Apply
$64k – $93k per year • Equity 0.1–0.5% • Remote/Hybrid • Full-Time • 6+ years exp • Paris
Node JS
Python
Rust
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Frontend
React.js
SvelteKit
DevOps
CloudFormation
Docker
GCP
Terraform
Apply
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 4 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
$221k – $387k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
DevOps
Incident Management
Management
ServiceNow
Apply
$56k – $94k per year • In office • Contractor • Santa Clara
Apply
$64k – $74k per year • In office • Contractor • Santa Clara
Apply
$56k per year • In office • Contractor • Santa Clara
Apply
$84k – $94k per year • In office • Contractor • Santa Clara
Apply
See all jobs
This is one of many
407,354 more open roles from verified company boards, updated every day.