Salary
≈ $44k – $106k per year (Estimated)
Location
In office (Shanghai, Beijing)
Seniority
Architect · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.
NVIDIA is leading company of AI computing. At NVIDIA, our employees are passionate about AI, HPC , VISUAL, GAMING. SA team is more focusing to bring NVIDIA new technology into difference industries. This role focuses on NVIDIA Inference Microservices (NIM), inference / RL rolloutperformance, and AI workflow enablement for LLM, VLM, and other generative AI workloads. It is a highly hands-on position at the intersection of model optimization, inference infrastructure, and customer solution delivery.
What you’ll be doing:
- Drive the implementation, deployment, and optimization of NVIDIA Inference Microservices (NIM) solutions for enterprise and industry AI workloads.
- Package and serve open-source, NVIDIA, and customer-proprietary models through NIM with standardized, containerized APIs for on-premises, cloud, and hybrid environments.
- Optimize high-volume inference and rollout workloads for LLMs and VLMs.
- Evaluate and tune the NIM models.
- Deliver technical projects, demos and client support tasks as directed by the Solution Architecture Leadership.
- Provide technical support and guidance to customers, facilitating the adoption and implementation of NVIDIA technologies and products.
- Collaborate with cross-functional teams to enhance and expand our AI solutions portfolio.
What we need to see:
- Master’s degree or higher in Computer Science, Machine Learning, Electrical Engineering, Mathematics, or a related technical field, or equivalent experience.
- 2+ years of hands-on experience in machine learning engineering, applied research, LLM/VLM inference, or RL rollout.
- Production-quality Python and PyTorch skills, including distributed GPU training, solution, profiling, debugging, memory optimization.
- Working knowledge of transformer architectures, performance optimization, rollout sampling strategies, structured generation, and model-quality evaluation.
- Strong written and verbal communication skills, with the ability to collaborate effectively across research, engineering, infrastructure, product, and customer-facing teams.
Ways to stand out from the crowd:
- Publications, open-source contributions, or significant technical projects, LLM/VLM, agent systems.
- Experience applying programmatic verification, simulators, compilers, execution sandboxes, APIs, or external tools as reward sources for model training. agent system.
- Familiar with oss RL framework such as SLIME, Nemo-RL.
- Familiarity with enterprise AI deployment, customer adaptation, or adapting foundation models to specialized vertical domains.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
697,651 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Shanghai
Principal AI Engineer (Platform team) - Improvado
10 hours ago
≈ $135k – $288k per year (Estimated) • Equity • Remote • 7+ years exp
Python
SQL
Databases
PostgreSQL
ClickHouse
RabbitMQ
AI/ML
Claude Code
AI Agents
LLM
LLM Guardrails
DevOps
Kubernetes
Analytics
ETL/ELT
Apply
Staff, Software Engineer
11 hours ago
≈ $129k – $242k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Sunnyvale
Python
Java
Databases
Azure Cosmos DB
AI/ML
AI Agents
LLM
Agentic Workflows
DevOps
Azure
CI/CD
Platform Engineering
Apply
Physics & Python Specialist
12 hours ago
Remote • Contractor • Master's Degree
Python
AI/ML
SciPy
NumPy
Gemini
LLM
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
Senior Networking Solutions Architect - Spectrum-X
10 hours ago
≈ $42k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Beijing • Shanghai • Shenzhen
Python
C++
DevOps
Linux
Apply
Senior System Engineer, Solution Engineering
10 hours ago
≈ $87k – $136k per year (Estimated) • Remote • Full-Time • 10+ years exp • Bachelor's Degree • Poland • Switzerland • Germany • Netherlands • Ukraine
AI/ML
InfiniBand
NVLink
DevOps
HPC
BGP
OSPF
Apply
Senior Solution Architect - ISV
10 hours ago
≈ $44k – $107k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
AI Agents
CUDA
DevOps
Platform Engineering
Apply
In office • Internship • Bachelor's Degree • Taipei • Hsinchu
Python
C
C++
Perl
C
Valgrind
DevOps
GitHub Actions
CircleCI
CI/CD
Jenkins
Git
Docker
Kubernetes
Spinnaker
KVM
QEMU
Xen
GitHub
GitLab
Apply
Senior Networking Solutions Architect - Spectrum-X
10 hours ago
≈ $42k – $101k per year (Estimated) • In office • Full-Time • 6+ years exp • Master's Degree • Beijing • Shanghai • Shenzhen
Python
C++
DevOps
Linux
Apply
Developer Relations Manager – CAE and CFD
10 hours ago
≈ $35k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • PhD • Shanghai • Beijing • Shenzhen
AI/ML
CUDA Toolkit
AI Agents
LLM
CUDA
Physical AI
Machine Learning
Apply
QNX- Director, Strategic Alliances
1 day ago
≈ $29k – $79k per year (Estimated) • In office • Full-Time • Shanghai
AI/ML
Physical AI
Apply
≈ $44k – $110k per year (Estimated) • In office • Full-Time • 2+ years exp • Shanghai • Beijing
C
C++
C
Pthreads
C++
TensorFlow C++
LLVM
AI/ML
CUDA Toolkit
Speech Recognition
TensorRT
OpenMP
TensorFlow
CUDA
cuDNN
MLIR
Apache TVM
CUTLASS
DevOps
CI/CD
Apply
Senior System Software Engineer, CPU
1 day ago
≈ $35k – $87k per year (Estimated) • In office • Full-Time • 5+ years exp • Shenzhen • Shanghai
C++
DevOps
Linux
Apply
≈ $36k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Beijing • Shanghai • Shenzhen
AI/ML
CUDA Toolkit
Reinforcement Learning
AI Agents
LLM
CUDA
Post-training
Physical AI
Machine Learning
Robotics
Reinforcement Learning
Apply
In office • Internship • PhD • Shanghai • Beijing • Shenzhen
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
JAX
SGLang
PyTorch
LLM
CUDA
Triton
NCCL
Chips/EDA
PoC Library
Apply
This is one of many
697,651 more open roles from verified company boards, updated every day.

