391,614open jobs
13,798companies
76,244added this week
Browse all
Salary
$179k – $372k per year (Estimated)
Location
In office (Zurich, United Kingdom, Poland, Germany)
Seniority
Architect · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are looking for a Senior GPU Networking Architect to join our networking software group, bringing strong GPU architecture and programming skills to build and improve GPU communication kernels. This role links GPU computing with networking by making sure communication primitives are carefully developed alongside GPU hardware capabilities. Join our team of engineers developing the software foundation for the largest AI systems globally.

What you will be doing:

  • Build, implement, and optimize GPU communication kernels that underpin collective and point-to-point operations in large-scale AI systems.

  • Leverage deep knowledge of GPU architecture-thread scheduling, memory hierarchy, execution pipelines-to improve kernel efficiency, minimize latency, and overlap computation with communication.

  • Develop GPU-resident communication primitives and device-side APIs that enable fine-grained, kernel-initiated data movement across nodes and accelerators.

  • Profile and tune GPU kernels end-to-end, identifying bottlenecks at the intersection of compute, memory, and network, and driving targeted optimizations.

  • Collaborate with network software, hardware, and AI framework teams to co-design communication strategies that align with GPU execution patterns and emerging model architectures.

  • Build proofs-of-concept, conduct experiments, and perform quantitative modeling to evaluate and validate new communication strategies before committing them to production.

  • Contribute to the evolution of programming models that expose GPU-aware networking capabilities to application developers.

What we need to see:

  • 5+ years of hands-on CUDA programming, including writing and optimizing non-trivial GPU kernels.

  • M.Sc. or equivalent experience in computer science, computer engineering, or a closely related field.

  • Strong understanding of GPU architecture fundamentals: warp scheduling, shared memory, L2 cache, memory coalescing, occupancy tuning, and asynchronous execution.

  • Experience with systems-level C/C++ development in performance-critical environments.

  • Familiarity with GPU data movement mechanisms such as GPUDirect RDMA and GPU-initiated communication.

  • Ability to read and reason about GPU performance profiles (e.g., Nsight Compute, Nsight Systems) and translate observations into actionable optimizations.

  • Strong collaboration skills in a multi-national, interdisciplinary environment.

Ways to stand out from the crowd:

  • Experience developing or optimizing communication kernels in libraries such as NCCL, NVSHMEM, or similar GPU-aware communication frameworks.

  • Understanding of distributed deep learning parallelism techniques, including data parallelism, tensor parallelism, pipeline parallelism, expert parallelism, and mixture-of-experts parallelism, and the communication patterns they impose on GPU kernels.

  • Background in RDMA, InfiniBand, high-speed networking, and GPU system topology, including NVLink, NVSwitch, PCIe, and network fabrics, and their impact on communication kernel design.

  • Experience with overlap techniques such as kernel pipelining, persistent kernels, or cooperative groups to hide communication latency behind compute.

  • Proven experience evaluating and optimizing large-scale LLM training or inference workloads, including hands-on work with frameworks such as PyTorch, TensorRT-LLM, or vLLM, and familiarity with emerging serving architectures such as disaggregated serving.

At NVIDIA, you'll work alongside colleagues who demonstrate deep expertise and innovative thinking in the industry, pushing the boundaries of what's possible in AI and high-performance computing. If you're passionate about GPU architecture, low-level kernel optimization, and building the communication fabric for next-generation AI, we want to hear from you!

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. For Poland: The base salary range is 292,500 PLN - 507,000 PLN for Level 4, and 375,000 PLN - 650,000 PLN for Level 5.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
391,614 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Zurich
$30k – $75k per year (Estimated) • Remote • 8+ years exp • Bachelor's Degree • Hyderabad
JavaScript
Python
SQL
TypeScript
AI/ML
AI Agents
LLM
Frontend
GraphQL
DevOps
Amazon EKS
Amazon S3
API Gateway
AWS
AWS Lambda
CI/CD
Docker
GitHub
Rest API
WebSockets
Kubernetes
Cybersecurity
Checkmarx
FedRAMP
ISO 27001
NIST 800-53
OWASP Top 10
Snyk
SOC 2
Threat Modeling
Veracode
Apply
Remote/Hybrid • 5+ years exp • Hyderabad
C#
Go
JavaScript
Python
TypeScript
AI/ML
AI Agents
LLM
DevOps
AWS
Azure
CI/CD
CloudFormation
GCP
GitHub
GitHub Actions
Helm
IAM
Jenkins
Kubernetes
Terraform
Cybersecurity
Checkmarx
Checkov
OWASP Top 10
Threat Modeling
Trivy
Veracode
Apply
Senior Developer 8 min ago
$18k – $60k per year (Estimated) • Remote/Hybrid • 6+ years exp • Hyderabad
C#
C++
JavaScript
Objective-C
Python
AI/ML
AI Agents
DevOps
CI/CD
Apply
$107k – $174k per year • In office • Full-Time • Bachelor's Degree • San Francisco
C++
Python
AI/ML
Computer Vision
Edge AI
Embodied AI
Multimodal AI
Reinforcement Learning
Self-Supervised Learning
Vision-Language-Action
Robotics
Imitation Learning
Motion Planning
Reinforcement Learning
ROS
ROS2
Sensor Fusion
SLAM
Apply
AI Analytics Lead 1 hour ago
$105k – $199k per year (Estimated) • Remote
SQL
Databases
ClickHouse
AI/ML
Anomaly Detection
LLM
LLM Guardrails
Analytics
A/B Testing
Metabase
Power BI
Tableau
Marketing
Amplitude
Mixpanel
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
Agentic Workflows
AI Agents
Function Calling
LangChain
LlamaIndex
LLM
OpenAI
Tool Use
DevOps
CI/CD
Git
Apply
In office • Internship • PhD • Shanghai
C++
Python
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenCL
DevOps
HPC
Apply
In office • Internship • Master's Degree • Shanghai
C#
C++
Python
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
AI Agents
Apply
In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
In office • Zurich
AI/ML
AI Agents
Management
Miro
Apply
In office • 3+ years exp • Zurich
Apply
Sales Director 2 days ago
$68k – $75k per year • In office • Full-Time • High School Diploma • Zurich
Apply
$78k – $184k per year (Estimated) • Remote/Hybrid • Full-Time • Paris • Palo Alto • New York • Zurich • Amsterdam
Python
AI/ML
DeepSpeed
Fine-tuning
FSDP
JAX
Knowledge Distillation
LLM
Megatron-LM
Mistral
Multimodal AI
Post-training
Pre-training
PyTorch
Ray
SFT
SkyPilot
Synthetic Data
Model Distillation
DevOps
Karpenter
Kubernetes
SLURM
Apply
Remote • Internship • 3+ years exp • PhD • Zurich
AI/ML
AI Agents
Marketing
Instagram
LinkedIn
Apply
See all jobs
This is one of many
391,614 more open roles from verified company boards, updated every day.