390,105open jobs
10,344companies
49,760added this week
Browse all
Salary
$272k – $431k per year
Location
In office (Santa Clara, United States, New York)
Seniority
Principal · 15+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is at the forefront of the AI revolution, and our research is shaping the future of large language models. We are looking for a Principal Scientist to set the technical direction for synthetic data generation across NVIDIA's frontier model efforts. You will define and build open-source libraries within the NVIDIA NeMo ecosystem that generate synthetic datasets across text, code, structured, and multimodal data, feeding the pre- and post-training of LLMs such as Nemotron. This role combines hands-on software engineering with applied research in generative methods, and you will collaborate with research, engineering, product, and model teams as well as external labs.

What you'll be doing:

  • Build and scale data generation pipelines using LLM-based methods combined with automated quality evaluation. resulting in datasets to improve both initial training and fine-tuning of LLMs such as Nemotron. These data pipelines cover reasoning, coding, structured output, and multimodal understanding.
  • Pioneer data generation for agentic and tool-use training: synthetic trajectories, multi-turn interactions, function calling, and executable environments for reinforcement learning, including reward modeling and verifiable-reward data
  • Advance multimodal synthetic data generation - image, document, video, and audio - in partnership with NVIDIA's model teams.
  • Advance privacy-preserving and safe synthesis - differential privacy, anonymization, and de-identification - enabling model training on sensitive data in regulated domains.
  • Develop and maintain open-source libraries and SDKs with clean APIs and strong documentation.
  • Drive software excellence with modern tooling, architecture based on configuration, and professional Git/CI-CD.
  • Publish original research at top machine learning and AI conferences to maintain NVIDIA's technical leadership.
  • Mentor scientists and engineers across the team, raising the technical bar and growing the next generation of researchers.

What we need to see:

  • PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.
  • 15+ years of engineering and research experience in synthetic data generation, generative modeling, multimodal machine learning, or related areas.
  • Deep technical understanding of LLMs, how data shapes their pre-training, post-training, and RL stages, and inference frameworks such as vLLM or TGI.
  • Proven track record of developing or maintaining software libraries used by a broad developer community.
  • Experience building and optimizing scalable data pipelines for large-scale model training - throughput, distributed inference, and cost at cluster scale.
  • Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL or similar.

Ways to stand out from the crowd:

  • Significant open-source contributions in ML or data tooling, with community adoption.
  • Experience with multimodal generation or understanding (vision-language, document AI, video, or audio).
  • Experience generating data for agentic, tool-use, or reinforcement-learning post-training, including RL environment design.
  • Background in differential privacy, de-identification, or synthetic data for regulated industries such as healthcare, finance, or government.
  • Experience influencing model training decisions at frontier scale, or partnering directly with pre-training and post-training teams.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and talented people in the world working with us. If you are creative, autonomous, and passionate about building open-source tools that make AI safer and more private, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 272,000 USD - 431,250 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until September 8, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
390,105 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
Data Architect 2 hours ago
In office • 10+ years exp • PhD
Python
SQL
Databases
Apache Kafka
Kafka
Snowflake
AI/ML
AI Agents
Anomaly Detection
Dagster
dbt
LangChain
LangGraph
LLM Guardrails
OpenAI
Prefect
RAG
DevOps
Amazon CloudWatch
Amazon Kinesis
Amazon S3
AWS
AWS Lambda
Azure
CI/CD
Cortex
Datadog
Vector
Prometheus
Analytics
ETL/ELT
Power BI
Tableau
Marketing
Salesforce
Apply
$165k – $306k per year • In office • Full-Time • 10+ years exp • San Antonio • Plano • Phoenix
AI/ML
AI Agents
Anthropic
AWS Bedrock
Context Engineering
Knowledge Graph
OpenAI
DevOps
AWS
Platform Engineering
Vector
Robotics
Path Planning
Apply
$127k – $172k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Springfield
PowerShell
Python
Mobile
JUnit
DevOps
CI/CD
GitLab
GitLab CI
Jenkins
Shift-Left
Cybersecurity
Nessus
OWASP ZAP
Shift-Left Security
QA
Cypress
Gatling
JMeter
Postman
Rest-Assured
Selenium
TestRail
Apply
Applied AI Engineer 2 hours ago
$152k – $271k per year • Remote/Hybrid • Full-Time • San Francisco • Austin • New York • Seattle
AI/ML
Fine-tuning
LLM
Multimodal AI
OCR
Prompt Engineering
RAG
Structured Outputs
Apply
$140k – $240k per year • Remote • Full-Time • Bachelor's Degree • United States
Java
Python
DevOps
Ansible
Azure
Chef
CI/CD
Configuration Management
Datadog
GCP
Grafana
Kubernetes
New Relic
Prometheus
Puppet
Splunk
Terraform
Cybersecurity
PCI DSS
SOC 2
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • PhD • Santa Clara
C++
Python
Apply
$144k – $230k per year • In office • Full-Time • PhD • Santa Clara
Apex
Apex
MuleSoft
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
In office • Full-Time • 5+ years exp • Master's Degree • Shanghai
C++
Python
AI/ML
AI Agents
Copilot
Cursor
LLM
Reinforcement Learning
Robotics
Isaac Lab
Isaac Sim
Perception
Reinforcement Learning
ROS
Sim-to-Real
Teleoperation
Apply
$73k – $175k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • São Paulo
AI/ML
CUDA
CUDA Toolkit
NLP
Synthetic Data
DevOps
Docker
HPC
Kubernetes
Apply
$104k – $212k per year (Estimated) • In office • Santa Clara
Apex
Apex
Lightning Web Components
Visualforce
DevOps
CI/CD
GitHub
Marketing
Salesforce
Apply
$104k – $212k per year (Estimated) • In office • Santa Clara
Apex
Apex
Lightning Web Components
Visualforce
DevOps
CI/CD
GitHub
Marketing
Salesforce
Apply
$145k – $294k per year (Estimated) • In office • Santa Clara
C++
C++
CMake
DevOps
Bazel
CI/CD
Docker
Platform Engineering
Apply
Salesforce Developer 4 hours ago
$93k – $211k per year (Estimated) • In office • Santa Clara
Apex
Go
Apex
Lightning Web Components
Visualforce
DevOps
CI/CD
Git
GitHub
Marketing
Salesforce
Apply
$200k – $322k per year • In office • Full-Time • Bachelor's Degree • Santa Clara
Go
Java
Python
JavaScript
AI/ML
AI Agents
Function Calling
LLM
Model Context Protocol
RAG
Frontend
React.js
DevOps
Incident Management
Management
Jira
ServiceNow
Apply
See all jobs
This is one of many
390,105 more open roles from verified company boards, updated every day.