657,845open jobs
38,355companies
93,289added this week
Browse all
Salary
$139k – $350k per year (Estimated)
Location
In office (Yokneam)
Seniority
Staff · 12+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is shaping the next era of computing, where AI, accelerated computing, and high-speed networking come together to power the world’s most advanced AI systems. Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect technologies - Ethernet, InfiniBand, NVLink, and BlueField DPUs - that connect thousands of GPUs into a single AI supercomputer.

NVIDIA is looking for a Technical Lead to join our Network System Validation group and lead the validation of advanced networking solutions across complex AI cluster environments. This is a deeply hands-on technical leadership role, combining ownership of the validation roadmap with technical mentoring and engineering excellence. You will develop validation methodologies and automation frameworks, while working hands-on on debugging, performance analysis, and cutting-edge AI networking technologies at scale. Join us to help push NVIDIA’s networking technologies to their limits and shape how next-generation AI infrastructure is validated.

What you’ll be doing:

  • Review system and product requirements, design validation methodologies, develop and implement comprehensive test plans, functional and performance, for networking technologies in large-scale AI cluster solutions
  • Develop and maintain benchmarks, automation tools and scripts for test execution, environment setup, log collection, and data analysis.
  • Lead end-to-end investigation of complex issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution
  • Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities
  • Collaborate deeply with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection
  • Profile and research AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations
  • Document findings, communicate technical results, and continuously improve validation methodologies, automation environments, and engineering processes

What we need to see:

  • B.Sc. / B.A. in Computer Science, Electrical Engineering, or equivalent experience
  • 12+ years of experience in networking, system validation, or related domains
  • Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause
  • Ability to read, debug, and reason about C/C++ code (Rust or Go a plus)
  • Strong scripting and automation experience using Python, Bash, and/or Ansible
  • Deep understanding of distributed systems: concurrency, consistency models, fault tolerance, and large-scale system performance under stress
  • Ability to drive technical alignment across teams, communicate tradeoffs clearly, and make high-quality architectural decisions at speed
  • Advance AI-driven approaches to test automation: intelligent scenario generation, LLM-augmented root-cause analysis, and autonomous validation pipelines

Ways to stand out from the crowd:

  • Experience with large-scale clusters or distributed systems
  • Familiarity with NVIDIA networking solutions (ConnectX, SpecX, BlueField)
  • Background in performance analysis, Kubernetes, or cloud environments
  • Background in chaos testing, fault injection, or simulation systems

We have some of the most forward-thinking and hardworking people working for us. If you're creative and autonomous, we want to hear from you! NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, disability status or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
657,845 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Yokneam
Senior Data Engineer 4 hours ago
$18k – $44k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Kuala Lumpur
Python
SQL
AI/ML
Claude
dbt
AI Agents
LLM
RAG
Hallucination
OpenAI
LLM Evaluation
LLM Guardrails
Agentic Workflows
Tool Use
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Platform Engineering
Vector
Cybersecurity
Least Privilege
Apply
$34k – $83k per year (Estimated) • In office • Internship • PhD • Shanghai
Python
Verilog
C++
Apply
$17k – $31k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
Reinforcement Learning
JAX
Computer Vision
TensorFlow
PyTorch
Synthetic Data
Physical AI
Robotics
Isaac Sim
MuJoCo
Isaac Lab
Imitation Learning
Reinforcement Learning
Apply
$168k – $270k per year • In office • Full-Time • 8+ years exp • PhD • Santa Clara
Python
C++
Fortran
AI/ML
CUDA Toolkit
CUDA
DevOps
HPC
Apply
$36k – $93k per year (Estimated) • In office • Internship • Shanghai
Python
Verilog
C++
SystemVerilog
Perl
AI/ML
NVLink
Chips/EDA
UVM
Apply
Senior NPI Engineer 4 hours ago
$131k – $237k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam
AI/ML
InfiniBand
Apply
$31k – $81k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Shanghai
AI/ML
Reinforcement Learning
TensorFlow
PyTorch
DevOps
GitHub
Robotics
Isaac Gym
Isaac Sim
MuJoCo
Isaac Lab
Motion Planning
Imitation Learning
Reinforcement Learning
Apply
$272k – $431k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Westford • Austin • Durham
DevOps
Platform Engineering
eBPF
IAM
HPC
Cybersecurity
SOC 2
Zero Trust
Apply
$152k – $230k per year • In office • Full-Time • 8+ years exp • PhD • Santa Clara
AI/ML
AI Agents
Apply
$144k – $342k per year (Estimated) • In office • Full-Time • 10+ years exp • Master's Degree • Tel Aviv
Apply
Senior NPI Engineer 4 hours ago
$131k – $237k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam
AI/ML
InfiniBand
Apply
$125k – $276k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam • Tel Aviv
AI/ML
ChatGPT
Edge AI
Apply
In office • Full-Time • 3+ years exp • PhD • Yokneam
Apply
$145k – $250k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Yokneam
AI/ML
InfiniBand
Apply
$86k – $196k per year (Estimated) • In office • Full-Time • 2+ years exp • Yokneam
Chips/EDA
Cadence Allegro
Apply
See all jobs
This is one of many
657,845 more open roles from verified company boards, updated every day.