429,021open jobs
14,503companies
63,779added this week
Browse all
Salary
$193k – $394k per year (Estimated)
Location
In office (Yokneam, Tel Aviv)
Seniority
Architect · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by excellent technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world.

NVIDIA is seeking a sharp, innovative, and hands-on Architect to help shape the future of LLM inference at scale. Join our dynamic E2E Architecture group, where we build ground breaking systems powering the next generation of generative AI workloads. In this role, you will work across software and hardware domains to design and optimize inference infrastructure for large language models running on some of the most sophisticated GPU clusters in the world. You’ll help define how AI models are deployed and scaled in production, driving decisions on everything from memory orchestration and compute scheduling to inter-node communication and system-level optimizations. This is an opportunity to work with top engineers, researchers, and partners across NVIDIA and leave a mark on the way generative AI reaches real-world applications.

What You’ll Be Doing:

  • Design and evolve scalable architectures for multi-node LLM inference across GPU clusters.

  • Develop infrastructure to optimize latency, throughput, and cost-efficiency of serving large models in production.

  • Collaborate with model, systems, compiler, and networking teams to ensure complete, high-performance solutions.

  • Prototype novel approaches to KV cache handling, tensor/pipeline parallel execution, and dynamic batching.

  • Evaluate and integrate new software and hardware technologies relevant to Core Spectrum-X technologies, such as load balancing, telemetry, congestion control, vertical application integration.

  • Work closely with internal teams and external partners to translate high-level architecture into reliable, high-performance systems.

  • Author design documents, internal specs, and technical blog posts and contribute to open-source efforts when appropriate.

What We Need to See:

  • Bachelor’s, Master’s, or PhD in Computer Science, Electrical Engineering, or equivalent experience.

  • 8+ years of experience building large-scale distributed systems or performance-critical software.

  • Deep understanding of deep learning systems, GPU acceleration, and AI model execution flows and/or high performance networking.

  • Proven software engineering skills in C++ and/or Python, preferably demonstrate strong familiarity with CUDA or similar platforms.

  • Strong system-level thinking across memory, networking, scheduling, and compute orchestration.

  • Excellent communication skills and ability to collaborate across diverse technical domains.

Ways to Stand Out from the Crowd:

  • Experience working on LLM - training or inference pipelines, transformer model optimization, or model-parallel deployments.

  • Proven success in profiling and optimizing performance bottlenecks across the LLM training or inference stack.

  • AI Accelerators and distributed communication patterns, congestion control and/or load balancing.

  • Shown optimization process for complex systems, deployed at scale to make impact.

  • Proven experience successfully driving complex organizational processes from planning through implementation.

NVIDIA is widely considered one of the most desirable places to work in tech - we are passionate about what we do and are committed to fostering a culture of perfection, innovation, and collaboration. If you’re excited to help define how the world runs AI at scale, this role is for you. Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. While planning your future, explore the benefits available to you and your family at www.nvidiabenefits.com/

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
429,021 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokneam
$61k – $171k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Yokneam
Python
Bash
Perl
AI/ML
CUDA Toolkit
CUDA
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Master's Degree • Santa Clara
Python
Java
C++
AI/ML
Claude
OpenAI Codex
DevOps
Prometheus
CI/CD
Apply
$152k – $242k per year • Remote • Full-Time • Bachelor's Degree • Redmond • Santa Clara
AI/ML
CUDA Toolkit
CUDA
Apply
$152k – $242k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
AI/ML
CUDA Toolkit
CUDA
MLIR
DevOps
HPC
Apply
In office • Internship • Master's Degree • Munich
Python
C++
DevOps
Git
Bazel
Apply
$184k – $288k per year • Remote • Full-Time • 6+ years exp • Bachelor's Degree • Austin
Python
DevOps
Puppet
Ansible
Chef
SLURM
CI/CD
GitOps
Configuration Management
Apply
$200k – $322k per year • In office • Full-Time • 12+ years exp • Master's Degree • Santa Clara
Python
Java
C++
AI/ML
Claude
OpenAI Codex
DevOps
Prometheus
CI/CD
Apply
$196k – $311k per year • In office • Full-Time • 12+ years exp • Master's Degree • Santa Clara
Python
Ruby
AI/ML
AI Agents
DevOps
GCP
Azure
AWS
IAM
Cybersecurity
MITRE ATT&CK
NIST CSF
Zero Trust
Threat Modeling
Apply
$152k – $242k per year • Remote • Full-Time • Bachelor's Degree • Redmond • Santa Clara
AI/ML
CUDA Toolkit
CUDA
Apply
$152k – $242k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
AI/ML
CUDA Toolkit
CUDA
MLIR
DevOps
HPC
Apply
$61k – $171k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Yokneam
Python
Bash
Perl
AI/ML
CUDA Toolkit
CUDA
Apply
$180k – $367k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
Apply
$112k – $277k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
Python
Java
AI/ML
Copilot
Model Context Protocol
AI Agents
NVLink
DevOps
CI/CD
AWS
Kubernetes
Platform Engineering
Gerrit
GitLab
Apply
$120k – $318k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
AI/ML
AI Agents
Apply
$61k – $171k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Yokneam
Python
Bash
Perl
AI/ML
CUDA Toolkit
CUDA
Apply
See all jobs
This is one of many
429,021 more open roles from verified company boards, updated every day.