433,445open jobs
14,773companies
64,942added this week
Browse all
Salary
$39k – $89k per year (Estimated)
Location
In office (Shenzhen)
Seniority
Architect · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

aNVIDIA’s Solutions Architect team is looking for a software-focused Solutions Architect to drive adoption of next-generation AI infrastructure across NVIDIA CPU platforms and LPU-based inference systems. This role will focus on NVIDIA CPUs, including Grace, Vera, and future CPU generations, and on LPU platforms and LPX-class systems used to accelerate large language model inference and other latency-sensitive generative AI workloads. We are looking for someone who understands that AI efficiency is a full-stack challenge spanning model architecture, runtime, compiler, serving framework, host software, memory movement, and workload partitioning across CPU, GPU, and LPU.

As a Solutions Architect, you will be the first line of technical expertise between NVIDIA and our customers for CPU- and LPU-centric AI system design. You will help customers understand how NVIDIA CPUs and LPU-based systems can improve the efficiency, latency, throughput, and total cost of their AI workloads, especially when deployed alongside NVIDIA GPUs in heterogeneous production environments. Your work will range from proof-of-concept development and software stack optimization to technical leadership with customer architects, engineering teams, and senior decision makers. You will engage directly with developers, ML engineers, researchers, platform architects, and IT leaders to identify bottlenecks, design optimization strategies, and build deployable reference architectures. You will also work closely with NVIDIA engineering, product, and field teams to translate customer needs into platform feedback, solution patterns, and roadmap inputs.

What you’ll be doing:

  • Evangelize NVIDIA CPU platforms, including Grace, Vera, and future generations, as well as LPU-based systems and LPX-class platforms, with a strong focus on AI software stacks and workload efficiency.

  • Help customers design and optimize AI workloads across CPU, GPU, and LPU, improving latency, throughput, utilization, and overall cost efficiency.

  • Analyze and tune LLM and generative AI pipelines across serving, runtime, memory, I/O, batching, scheduling, and orchestration layers.

  • Build proof-of-concepts, reference architectures, and technical guidance in partnership with Engineering, Product, and Sales teams.

  • Establish trusted technical relationships with customer architects, infrastructure teams, and senior leaders, becoming a strategic advisor for heterogeneous AI system design.

What we need to see:

  • MS or PhD in Computer Science, Engineering, Mathematics, Physics, or a related field, or equivalent experience, plus 5+ years in AI systems, infrastructure, performance engineering, or solution architecture.

  • Strong understanding of modern CPU architecture, Linux systems, and software performance tuning, along with hands-on experience in AI inference for LLM, generative AI, or agentic AI workloads.

  • Experience optimizing heterogeneous systems involving CPU and accelerators, with familiarity in frameworks such as PyTorch, Triton, TensorRT-LLM, vLLM, or ONNX Runtime.

  • Strong programming, problem-solving, and communication skills, with the ability to work effectively with both technical teams and senior customer stakeholders.

Ways to stand out from the crowd:

  • Experience with NVIDIA CPU platforms such as Grace, Grace Hopper, or Arm64 server environments, and familiarity with LPU-based systems or other low-latency inference accelerators.

  • Deep expertise in LLM inference optimization, serving architecture, and workload placement across CPU, GPU, and LPU.

  • Experience building customer-facing proof-of-concepts and measuring AI efficiency through latency, throughput, cost per token, power, or utilization.

  • Familiarity with NVIDIA AI software and platform technologies.

NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-looking and talented people in the world working with us. If you are creative, autonomous, and excited about helping customers build highly efficient AI platforms across CPU, GPU, and LPU technologies, we want to hear from you.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. We highly value diversity in our current and future employees and do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
433,445 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shenzhen
$22k per year (net) • Remote • Minsk
Go
JavaScript
PHP
TypeScript
SQL
PHP
Symfony
Databases
PostgreSQL
Redis
AI/ML
Copilot
Cursor
Claude
ChatGPT
Claude Code
Embeddings
Prompt Engineering
AI Agents
Gemini
LLM
RAG
OpenAI
Anthropic
Frontend
React.js
DevOps
Rest API
GCP
GitHub Actions
WebSockets
CI/CD
Docker
Kubernetes
Google GKE
GitHub
Apply
$126k – $280k per year (Estimated) • Remote/Hybrid • London
AI/ML
AI Agents
Apply
$83k – $200k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Barcelona
AI/ML
AI Agents
Agentic Workflows
Apply
In office • Full-Time • Bachelor's Degree • Mumbai
AI/ML
AI Agents
DevOps
SLI/SLO/SLA
Apply
$74k – $164k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Singapore
AI/ML
AI Agents
Agentic Workflows
Apply
$14k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
LangChain
LlamaIndex
Function Calling
AI Agents
LLM
OpenAI
Agentic Workflows
Tool Use
DevOps
CI/CD
Git
Apply
$15k – $24k per year (Estimated) • In office • Internship • PhD • Shanghai
Python
C
C++
C
MPI
AI/ML
CUDA Toolkit
OpenCL
CUDA
DevOps
HPC
Apply
$14k – $22k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
C#
C++
AI/ML
Model Context Protocol
AI Agents
Apply
$27k – $68k per year (Estimated) • In office • Internship • Master's Degree • Shanghai
Python
Perl
AI/ML
AI Agents
Apply
$14k – $22k per year (Estimated) • In office • Internship • PhD • Shanghai
AI/ML
AI Agents
RAG
Apply
$17k – $47k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Shenzhen
Apply
In office • Shenzhen
Apply
$20k – $56k per year (Estimated) • Remote • Full-Time • Shenzhen
Apply
Product Manager 2 days ago
$24k – $57k per year (Estimated) • Remote • Full-Time • Shenzhen
Apply
$32k – $67k per year (Estimated) • In office • Contractor • Shenzhen
DevOps
Amazon S3
Apply
See all jobs
This is one of many
433,445 more open roles from verified company boards, updated every day.