402,911open jobs
14,044companies
78,108added this week
Browse all
Location
In office (Shanghai, Beijing, Shenzhen)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking a passionate, world-class software engineer to join its Compute Developer Technology team(DevTech). Our team has over 150 engineers across Beijing, Shanghai, Shenzhen, Taipei, Seoul, and Sydney. We understand algorithms, GPU, and real-world applications. Our mission is to connect the NVIDIA platform with developers worldwide. We dive deep into customer projects to solve performance bottlenecks. We use insights from workloads to guide next-generation NVIDIA hardware and software. If you are driven by innovation and ambition, this is the team for you!

What you'll be doing:

  • Working directly with key application developers to understand the current and future problems they are solving. You will build and optimize core parallel algorithms and data structures to deliver the most effective solutions using GPUs, through both library development and direct contribution to applications. This includes training and inference optimization for large language models (LLM), contributing to frameworks and open-source projects in the large language models ecosystem, such as Megatron and TRTLLM, SGLang, vLLM...

  • Collaborating closely with the architecture, research, libraries, tools, and system software teams at NVIDIA to influence the build of next-generation architectures, software platforms, and programming models. This includes investigating impact on application performance and developer efficiency, and turning real-world developer feedback into actionable platform improvements.

  • Engaging in deep optimization of high-performance operators, involving but not limited to GPU kernel optimization, instruction-level tuning, and compiler optimization. These optimizations will directly support customers or be coordinated within computation libraries and open-source projects across the community, like cuDNN, cuBLAS, and CUTLASS and Open- source libs like DeepGEMM, FlashMLA, FlashAttention, Flashinfer...

  • Improving communication for broad distributed large language models workloads. You will spearhead advancements in distributed training and inference by refining communication libraries(NCCL,NCCL GIN , NVSHMEM) and engaging in open-source communication libraries(like DeepEP, NCCL EP). This demands in-depth study of interconnect topologies(NVLINK) and network protocols(InfiniBand/RoCE) to design efficient data transfer strategies and methods for compute-communication overlap.

What we need to see:

  • A degree or equivalent experience from a university in an engineering or computer science related field. A masters or doctoral degree is preferred.

  • 2+ years of work experience.

  • Solid understanding of C, C++, Python, or Fortran.

  • Strong knowledge of software development, programming techniques, and algorithms.

  • Strong mathematical fundamentals, including linear algebra and numerical methods.

  • Background in parallel programming and accelerated computing, with comprehensive knowledge of parallel architectures and methods for performance analysis and tuning. Experience in GPU programming is desirable.

  • Experience in full-stack performance analysis and optimization within at least one of these areas: large language models and high-performance computing. Having expertise ranging from operator-level through framework-level to algorithm-level optimization is strongly preferred.

  • Experience in distributed communication optimization is highly advantageous. This involves familiarity with remote direct memory access, GPU interconnects, collective communication algorithms, and associated open-source libraries used in large-scale model training and inference.

  • Solid software engineering fundamentals and system architecture thinking, with the ability to build modules and drive engineering practices in complex systems.

  • Strong communication and cooperation abilities, with the capability to work efficiently alongside architecture, research, and software product teams to promote optimization from concept to production.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
402,911 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$5.3k – $16k per year • In office • Full-Time • Gurgaon
JavaScript
Python
TypeScript
AI/ML
AI Agents
Claude
Claude Code
LLM
Marketing
HubSpot
LinkedIn
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$48k – $52k per year • Equity • In office • Bachelor's Degree
C#
C++
Python
Verilog
VHDL
Apply
$48k – $52k per year • Equity • In office • Bachelor's Degree
C#
C++
JavaScript
TypeScript
C#
.NET
Frontend
Angular
DevOps
Azure
Marketing
Salesforce
Apply
Fullstack Engineer 2 hours ago
$100k – $150k per year • Equity 0.2–2% • Remote • Full-Time • 1+ year exp • San Francisco
C++
JavaScript
Python
DevOps
AWS
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
Agentic Workflows
AI Agents
Function Calling
LangChain
LlamaIndex
LLM
OpenAI
Tool Use
DevOps
CI/CD
Git
Apply
In office • Internship • PhD • Shanghai
C++
Python
C
C
MPI
AI/ML
CUDA
CUDA Toolkit
OpenCL
DevOps
HPC
Apply
In office • Internship • Master's Degree • Shanghai
C#
C++
Python
AI/ML
AI Agents
Model Context Protocol
Apply
In office • Internship • Master's Degree • Shanghai
Perl
Python
AI/ML
AI Agents
Apply
In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
In office • 3+ years exp • Bachelor's Degree • Shanghai
Apply
In office • Full-Time • 10+ years exp • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
ABAP
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
Apply
In office • Internship • 8+ years exp • Shanghai
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
See all jobs
This is one of many
402,911 more open roles from verified company boards, updated every day.