430,068open jobs
14,667companies
59,611added this week
Browse all
Salary
$118k – $343k per year (Estimated)
Location
In office (Yokneam, Israel)
Seniority
Staff · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect - Ethernet, InfiniBand, NVLink, and BlueField DPUs - that switches thousands of GPUs into a single AI supercomputer, moving data at the scale and speed the most demanding workloads require. NVIDIA is looking for an Engineering Manager to join our Network System Validation group. You will work on system validating advanced networking solutions across NVIDIA complex AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove our network's correctness and performance at scale. In this role you will lead the validation direction and engineering excellence of one of our technology validation teams. This is a management role for a technology leader who can own the technical roadmap and execution, mentoring a team of high-performance engineers, and push NVIDIA's network to its speed-of-light limits. This role combines the development of methodologies and automation tools with system validation, performance analysis, and investigation of cutting-edge AI networking technologies at scale.

What you’ll be doing:

  • Lead, mentor, and coach a team of software development and system validation engineers.

  • Review system and product requirements, design validation methodologies, develop comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions

  • Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.

  • Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution

  • Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection

  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations

  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes

  • Foster a team culture centered on software quality, accountability, and technical excellence

What we need to see:

  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience

  • 8+ overall years of experience in networking, system validation, or related domains

  • 3+ years of experience leading software or system development team

  • Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause

  • Strong scripting and automation experience using Python, Bash, and/or Ansible

  • Ability to read, debug, and reason about C/C++ code (Rust or Go a plus)

  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed

  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines

Ways to stand out from the crowd:

  • Experience with large-scale clusters or distributed systems

  • Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)

  • Background in performance analysis, Kubernetes, or cloud environments

  • Background in chaos testing, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you! NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, disability status or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
430,068 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokneam
$36k – $96k per year (Estimated) • Remote • Full-Time • 6+ years exp
Rust
C++
DevOps
Splunk
Terraform
CloudFormation
Envoy
HAProxy
AWS
Docker
eBPF
SLI/SLO/SLA
Amazon CloudWatch
Apply
$24k – $56k per year (Estimated) • Remote • Full-Time • 6+ years exp
Rust
C++
DevOps
Splunk
Terraform
CloudFormation
Envoy
HAProxy
AWS
Docker
eBPF
SLI/SLO/SLA
Amazon CloudWatch
Apply
$34k – $92k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Python
SQL
Python
FastAPI
Django
Celery
Databases
PostgreSQL
Redis
AI/ML
NLP
LLM
RAG
DevOps
Rest API
gRPC
WebSockets
CI/CD
Docker
Kubernetes
Analytics
ETL/ELT
Apply
Software Developer 1 hour ago
$37k – $98k per year (Estimated) • In office • Full-Time • Tokyo
Python
DevOps
Docker
Kubernetes
Apply
$128k – $312k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • Sydney
JavaScript
TypeScript
C++
Frontend
React.js
Design
Canva
Apply
$143k – $295k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • Yokneam
AI/ML
InfiniBand
DevOps
GitHub
Analytics
Power BI
Microsoft Excel
Management
Confluence
Jira
Apply
Research Scientist 1 hour ago
In office • Full-Time • 2+ years exp • PhD • Singapore
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
Computer Vision
PyTorch
LLM
CUDA
Apply
$23k – $52k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Shanghai
Analytics
Microsoft Excel
Apply
$184k – $288k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara • Austin • Hillsboro • Boulder • Redmond
AI/ML
CUDA Toolkit
CUDA
DevOps
HPC
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara • Redmond • Seattle
AI/ML
Computer Vision
Synthetic Data
World Models
Robotics
Sensor Fusion
Apply
$61k – $171k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Yokneam
Python
Bash
Perl
AI/ML
CUDA Toolkit
CUDA
Apply
$193k – $394k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
Python
C++
AI/ML
CUDA Toolkit
LLM
CUDA
KV Cache
Apply
$180k – $367k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
Apply
$112k – $277k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Tel Aviv • Yokneam
Python
Java
AI/ML
Copilot
Model Context Protocol
AI Agents
NVLink
DevOps
CI/CD
AWS
Kubernetes
Platform Engineering
Gerrit
GitLab
Apply
$120k – $318k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
AI/ML
AI Agents
Apply
See all jobs
This is one of many
430,068 more open roles from verified company boards, updated every day.