412,170open jobs
14,281companies
71,440added this week
Browse all
Salary
$43k – $97k per year (Estimated)
Location
In office (Shanghai, Beijing)
Seniority
Architect
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking Software Performance Architects to optimize GPU kernel performance for state-of-the-art data-center platforms. We build automated, data-driven workflows to detect, explain, and prevent performance regressions across key deep learning workloads, partnering closely with kernel developers, compiler teams, infrastructure, and architecture/performance groups.

What you'll be doing:

  • Performance analysis, optimization and debugging

    • Build performance narratives using structured methodology: baselines, projections, controlled comparisons, and regression attribution.

    • With the methodologies, analyze performance of GPU-accelerated kernels and key deep learning building blocks, identify gaps with baselines or projections, then optimize the kernels' performance to fill the gaps.

    • Debug performance issues end-to-end: reproduce, isolate root causes, propose fixes or mitigation paths, and drive closure with the owning teams.

  • Automation + regression infrastructure (Python-heavy)

    • Develop and maintain Python-based automation for performance testing and analysis-using modern AI-assisted developer tools (e.g., Cursor/Claude Code/Copilot) to accelerate scripting while keeping code maintainable and reviewable.

    • Design and operate performance test workflows: coverage definition, test/workload generation, automated large-scale execution (CI/nightly/on-demand), rerun rules, and reproducibility standards.

  • Cross-team collaboration and operating model

    • Work with kernel developers and the compiler teams to ensure performance checks are practical, scalable, and aligned to release needs.

    • Work with chip architecture and modeling teams to solidify the performance methodology across chip architecture generations and common Deep Learning operators such as GEMM, Attention, MoE.

    • Partner with SWQA and infrastructure teams for execution at scale and reliable pipelines/dashboards.

  • Following general software engineering best practices including support for regression testing and CI/CD flows

What we need to see:

  • Masters or PhD degree or equivalent experience in Computer Science, Computer Engineering, Applied Math, or related field

  • Strong programming ability in Python plus C/C++ with 2+ working experience (performance-oriented code reading/debugging)

  • Solid fundamentals in computer architecture, parallel programming and performance reasoning (latency/throughput, memory hierarchy, parallelism) to be able to identify bottlenecks, optimize resource utilization, and improve throughput

  • Experience with performance analysis workflows: profiling, measurement methodology, reproducibility, and regression triage.

  • Comfortable working across teams and driving issues to decision/closure with clear communication

Ways to stand out from the crowd:

  • Experience with high-performance kernels or math libraries (e.g., GEMM/attention, CUTLASS-like concepts)

  • GPU programming/perf experience (CUDA or equivalent parallel programming)

  • Strong ML/DL workload understanding (training/inference shapes, precision modes, perf bottlenecks)

  • Familiarity with simulators/analytical modeling or performance characterization methodology

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
412,170 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
Principle Engineer 25 min ago
Remote/Hybrid • Full-Time • Bengaluru
Python
Java
C++
Management
Telegram
WhatsApp
Apply
$162k – $222k per year • Remote/Hybrid • Secret • 10+ years exp • Bachelor's Degree
Python
Java
C++
AI/ML
Human-in-the-Loop
DevOps
Debian
CI/CD
Git
Robotics
Sensor Fusion
Apply
$20k – $52k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
Python
C++
Bash
C++
CMake
DevOps
CI/CD
Git
QEMU
GitLab
Management
Confluence
Apply
$150k – $175k per year • Equity • In office • 5+ years exp
AI/ML
Claude
Design
Figma
Apply
$33k – $40k per year (net) • In office • Moscow
Python
PHP
Bash
Databases
PostgreSQL
DevOps
PHP-FPM
Terraform
Ansible
Zabbix
Loki
Prometheus
GitLab CI
CI/CD
Jenkins
Git
Docker
Kubernetes
Nginx
Grafana
CentOS Stream
GitLab
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
LangChain
LlamaIndex
Function Calling
AI Agents
LLM
OpenAI
Agentic Workflows
Tool Use
DevOps
CI/CD
Git
Apply
$14k – $23k per year (Estimated) • In office • Internship • PhD • Shanghai
Python
C
C++
C
MPI
AI/ML
CUDA Toolkit
OpenCL
CUDA
DevOps
HPC
Apply
$13k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
C#
C++
AI/ML
Model Context Protocol
AI Agents
Apply
In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
AI Agents
Apply
$13k – $21k per year (Estimated) • In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$20k – $51k per year (Estimated) • Remote • Full-Time • 6+ years exp • High School Diploma • Shanghai
Apply
$18k – $45k per year (Estimated) • In office • Full-Time • Master's Degree • Beijing • Shanghai
Apply
$19k – $44k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree • Guangzhou • Zhongshan • Zhaoqing • Shantou • Zhanjiang
Cybersecurity
CAPA
Apply
$18k – $39k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Shanghai
Apply
$15k – $32k per year (Estimated) • In office • Full-Time • 1+ year exp • Associate's Degree • Shanghai • Wuxi
Apply
See all jobs
This is one of many
412,170 more open roles from verified company boards, updated every day.