399,808open jobs
13,937companies
77,696added this week
Browse all
Location
In office (Shanghai, Beijing)
Seniority
Junior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology-and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.

We are now looking for cuTile Core Compiler Architect in our group! The NVIDIA Architecture group is looking for world class architects and engineers to join and lead our various architecture efforts. A key part of NVIDIA's strength is to innovate in the graphics and parallel computing fields delivering the highest performance in the world for parallel processing algorithms. We are constantly looking for ways to improve our GPU architecture and maintain our leadership by developing new parallel programming models, new architectures and new infrastructure that is required to make this successful.

What you'll be doing:

  • Design and implement the DSL and the core compiler of tile-aware GPU programming model for emerging GPU architectures

  • Continuously innovate and iterate on the core architecture of the compiler to consistently optimize performance

  • Investigation of next-generation GPU architectures and provide solutions in the DSL and compiler stack

  • Performance analysis on emerging AI/LLM workloads and integrate with AI/ML frameworks

What we need to see:

  • Masters or PhD or equivalent experience in relevant discipline (CE, CS&E, CS, AI)

  • 2+ years of relevant work experience

  • Excellent C/C++ programming and software engineering skills, ACM background is a plus

  • Good fundamental knowledges on computer architecture

  • Strong ability in abstracting problems and the methodology in resolving problems

  • Strong compiler backgrounds including MLIR/TVM/Triton/LLVM is desired

  • Good knowledge of GPU architecture and fast kernel programming skills is a plus

  • Knowledge of LLM algorithms or a certain HPC domain is a plus

  • Knowledge of multi-GPU distributed communication is a plus

  • Excellent oral communication in English is a plus

Widely considered to be one of the technology world’s most desirable employers, NVIDIA offers highly competitive salaries and a comprehensive benefits package. As you plan your future, see what we can offer to you and your family www.nvidiabenefits.com/

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
399,808 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$90k – $171k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
LangChain
Langfuse
LLM
OpenAI
Structured Outputs
Tool Use
Management
Slack
Apply
$56k – $128k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
LangChain
Langfuse
LLM
OpenAI
Structured Outputs
Tool Use
Management
Slack
Apply
$72k – $149k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
LangChain
Langfuse
LLM
OpenAI
Structured Outputs
Tool Use
Management
Slack
Apply
$52k – $100k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
LangChain
Langfuse
LLM
OpenAI
Structured Outputs
Tool Use
Management
Slack
Apply
$101k – $189k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp
AI/ML
AI Agents
Anthropic
Function Calling
Human-in-the-Loop
LangChain
Langfuse
LLM
OpenAI
Structured Outputs
Tool Use
Management
Slack
Apply
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 3 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
In office • 3+ years exp • Bachelor's Degree • Shanghai
Apply
In office • Full-Time • 10+ years exp • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
ABAP
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
Apply
In office • Internship • 8+ years exp • Shanghai
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
See all jobs
This is one of many
399,808 more open roles from verified company boards, updated every day.