697,651open jobs
41,257companies
99,088added this week
Browse all
Salary
$44k – $106k per year (Estimated)
Location
In office (Shanghai, Beijing)
Seniority
Architect · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. SA team is more focusing to bring NVIDIA new technology into difference industries. This role focuses on NVIDIA Inference Microservices (NIM), inference / RL rolloutperformance, and AI workflow enablement for LLM, VLM, and other generative AI workloads. It is a highly hands-on position at the intersection of model optimization, inference infrastructure, and customer solution delivery.

What you’ll be doing:

  • Drive the implementation, deployment, and optimization of NVIDIA Inference Microservices (NIM) solutions for enterprise and industry AI workloads.
  • Package and serve open-source, NVIDIA, and customer-proprietary models through NIM with standardized, containerized APIs for on-premises, cloud, and hybrid environments.
  • Optimize high-volume inference and rollout workloads for LLMs and VLMs.
  • Evaluate and tune the NIM models.
  • Deliver technical projects, demos and client support tasks as directed by the Solution Architecture Leadership.
  • Provide technical support and guidance to customers, facilitating the adoption and implementation of NVIDIA technologies and products.
  • Collaborate with cross-functional teams to enhance and expand our AI solutions portfolio.

What we need to see:

  • Master’s degree or higher in Computer Science, Machine Learning, Electrical Engineering, Mathematics, or a related technical field, or equivalent experience.
  • 2+ years of hands-on experience in machine learning engineering, applied research, LLM/VLM inference, or RL rollout.
  • Production-quality Python and PyTorch skills, including distributed GPU training, solution, profiling, debugging, memory optimization.
  • Working knowledge of transformer architectures, performance optimization, rollout sampling strategies, structured generation, and model-quality evaluation.
  • Strong written and verbal communication skills, with the ability to collaborate effectively across research, engineering, infrastructure, product, and customer-facing teams.

Ways to stand out from the crowd:

  • Publications, open-source contributions, or significant technical projects, LLM/VLM, agent systems.
  • Experience applying programmatic verification, simulators, compilers, execution sandboxes, APIs, or external tools as reward sources for model training. agent system.
  • Familiar with oss RL framework such as SLIME, Nemo-RL.
  • Familiarity with enterprise AI deployment, customer adaptation, or adapting foundation models to specialized vertical domains.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
697,651 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$135k – $288k per year (Estimated) • Equity • Remote • 7+ years exp
Python
SQL
Databases
PostgreSQL
ClickHouse
RabbitMQ
AI/ML
Claude Code
AI Agents
LLM
LLM Guardrails
DevOps
Kubernetes
Analytics
ETL/ELT
Apply
$129k – $242k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Sunnyvale
Python
Java
Databases
Azure Cosmos DB
AI/ML
AI Agents
LLM
Agentic Workflows
DevOps
Azure
CI/CD
Platform Engineering
Apply
Remote • Contractor • Master's Degree
Python
AI/ML
SciPy
NumPy
Gemini
LLM
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
$42k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Beijing • Shanghai • Shenzhen
Python
C++
DevOps
Linux
Apply
$87k – $136k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • Poland • Switzerland • Germany • Netherlands • Ukraine
AI/ML
InfiniBand
NVLink
DevOps
HPC
BGP
OSPF
Apply
$44k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
AI Agents
CUDA
DevOps
Platform Engineering
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
$42k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Beijing • Shanghai • Shenzhen
Python
C++
DevOps
Linux
Apply
$35k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • PhD • Shanghai • Beijing • Shenzhen
AI/ML
CUDA Toolkit
AI Agents
LLM
CUDA
Physical AI
Machine Learning
Apply
$29k – $79k per year (Estimated) • In office • Full-Time • Shanghai
AI/ML
Physical AI
Apply
$44k – $110k per year (Estimated) • In office • Full-Time • 2+ years exp • Shanghai • Beijing
C
C++
C
Pthreads
C++
TensorFlow C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
TensorRT
OpenMP
TensorFlow
CUDA
cuDNN
MLIR
Apache TVM
CUTLASS
DevOps
CI/CD
Apply
$35k – $87k per year (Estimated) • In office • Full-Time • 5+ years exp • Shenzhen • Shanghai
C++
DevOps
Linux
Apply
$36k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Beijing • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
Reinforcement Learning
AI Agents
LLM
CUDA
Post-training
Physical AI
Machine Learning
Robotics
Reinforcement Learning
Apply
In office • Internship • PhD • Shanghai • Beijing • Shenzhen
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
Chips/EDA
PoC Library
Apply
See all jobs
This is one of many
697,651 more open roles from verified company boards, updated every day.