406,988open jobs
14,118companies
78,699added this week
Browse all
Salary
$136k – $343k per year (Estimated)
Location
Remote/Hybrid (Tel Aviv, Israel)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been at the forefront of the deep learning revolution, pioneering innovations that have transformed the entire field. As the leading provider of GPUs and AI computing platforms, NVIDIA has empowered researchers and engineers worldwide to accelerate breakthroughs in artificial intelligence.

We seek a versatile Senior Software Engineer who is passionate about performance optimization and generative AI. Our team brings the latest research in LLM inference - from novel decoding strategies to quantization schemes - into production across NVIDIA's hardware lineup, from large data center servers to powerful edge devices. We work on the most advanced architectures in the field, with a focus on NVIDIA's own.

What you'll be doing:

  • Implement and optimize inference algorithms for LLM and omnimodal architectures, including hybrid Mamba-Transformer and mixture-of-experts models

  • Profile inference pipelines using NVIDIA's profiling and simulation tools. Correlate simulation predictions against real hardware across data center and edge devices

  • Write and tune GPU kernels (CUDA, Triton) for operators like fused MoE layers, SSM state updates, and quantized GEMMs

  • Solve distributed inference problems: expert parallelism, communication-compute overlap, collective tuning, multi-node deployment

  • Build production-grade software inside major open-source libraries - vLLM, SGLang, Dynamo, FlashInfer

  • Own optimization features end-to-end, from scoping through delivery, collaborating with research, product, and engineering teams worldwide

What we need to see:

  • B.Sc., M.Sc., or equivalent experience in Computer Science or Computer Engineering

  • 5+ years of hands-on software engineering experience in performance-critical systems

  • Solid understanding of deep learning architectures (Transformers, SSMs, MoE, …)

  • Experience with systems where hardware constraints matter: GPU programming, memory hierarchy, networking, or distributed computing

  • Strong software engineering fundamentals: clean design, extensibility, testability. Good judgment about when complexity is warranted

  • Effective communicator who works well across teams and time zones

  • Experience optimizing deep learning workloads on NVIDIA GPUs using roofline models, Nsight/PyTorch profilers and end-to-end traces

Ways to stand out from the crowd:

  • Contributions to open-source inference runtimes and libraries - vLLM, SGLang, FlashInfer, Dynamo or similar

  • Hands-on work with LLM quantization (FP8, NVFP4, MXFP8, mixed-precision) and practical understanding of numerical precision tradeoffs

  • Track record with distributed inference at scale: tensor parallelism, pipeline parallelism, expert parallelism, disaggregation, multi-node orchestration

  • Deep knowledge of the latest LLM architectural trends: multi-token predictors, sparse hybrid models, attention and state-space mechanisms

  • Experience with performance modeling and simulation-to-silicon correlation

NVIDIA is widely considered one of the world's most desirable employers in the technology field. We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you! We are committed to fostering a diverse work environment and are proud to be an equal-opportunity employer. We highly value diversity in our current and future employees. We do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
406,988 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Tel Aviv
Product Engineer 1 day ago
$150k – $225k per year • Equity 0.2–0.4% • Remote • Full-Time • 3+ years exp • San Francisco
Python
Python
FastAPI
AI/ML
AI Agents
LLM
Frontend
Next.js
DevOps
Azure
WebSockets
Apply
$200k – $260k per year • Remote • Full-Time • 10+ years exp
JavaScript
Node JS
TypeScript
Node JS
Electron
AI/ML
LLM
Apply
$113k – $247k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
Anthropic
Claude
Claude Code
Context Engineering
Embeddings
Function Calling
Hallucination
Hugging Face
Human-in-the-Loop
LangChain
LangGraph
LlamaIndex
LLM
LLM Guardrails
LLMOps
Model Context Protocol
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
RAG
Reranking
Scikit-learn
Semantic Kernel
Spark
Structured Outputs
TensorFlow
Vertex AI
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Amazon S3
AWS
AWS Lambda
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
Analytics
ETL/ELT
Matplotlib
Power BI
Seaborn
Tableau
Management
ServiceNow
Apply
$73k – $160k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
Anthropic
Claude
Claude Code
Context Engineering
Embeddings
Function Calling
Hallucination
Hugging Face
Human-in-the-Loop
LangChain
LangGraph
LlamaIndex
LLM
LLM Guardrails
LLMOps
Model Context Protocol
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
RAG
Reranking
Scikit-learn
Semantic Kernel
Spark
Structured Outputs
TensorFlow
Vertex AI
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Amazon S3
AWS
AWS Lambda
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
Analytics
ETL/ELT
Matplotlib
Power BI
Seaborn
Tableau
Management
ServiceNow
Apply
$96k – $211k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
AI Agents
Amazon SageMaker
Anomaly Detection
Anthropic
Claude
Claude Code
Context Engineering
Embeddings
Function Calling
Hallucination
Hugging Face
Human-in-the-Loop
LangChain
LangGraph
LlamaIndex
LLM
LLM Guardrails
LLMOps
Model Context Protocol
NLP
NumPy
Pandas
Prompt Engineering
PyTorch
RAG
Reranking
Scikit-learn
Semantic Kernel
Spark
Structured Outputs
TensorFlow
Vertex AI
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Amazon S3
AWS
AWS Lambda
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
Analytics
ETL/ELT
Matplotlib
Power BI
Seaborn
Tableau
Management
ServiceNow
Apply
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 4 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
In office • 3+ years exp • Tel Aviv
Apply
IT Support Engineer 9 hours ago
$60k – $131k per year (Estimated) • In office • Full-Time • 3+ years exp • Tel Aviv
DevOps
Ansible
Azure
Management
Jira
Slack
Apply
Game Engineer 10 hours ago
$29k – $111k per year (Estimated) • In office • Full-Time • 4+ years exp • Tel Aviv
C#
Go
Node JS
JavaScript
DevOps
AWS
Azure
CI/CD
GCP
WebSockets
Game Dev
Unity
Analytics
A/B Testing
Apply
SecOps Engineer 10 hours ago
Remote/Hybrid • Full-Time • Tel Aviv
Python
AI/ML
Anomaly Detection
DevOps
AWS
Azure
GCP
SLI/SLO/SLA
Splunk
Apply
$81k – $231k per year (Estimated) • In office • 5+ years exp • Tel Aviv
C#
C++
Go
Java
Kotlin
Scala
TypeScript
JavaScript
Databases
PostgreSQL
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
See all jobs
This is one of many
406,988 more open roles from verified company boards, updated every day.