397,589open jobs
13,889companies
77,448added this week
Browse all
Location
In office (Shanghai, Beijing, Shenzhen)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
NVIDIA is an American technology company founded in 1993 that invented the graphics processing unit and has become the dominant supplier of accelerated computing platforms for artificial intelligence. Its portfolio spans data centre GPUs and systems built on the Hopper and Blackwell architectures, GeForce consumer graphics, automotive and robotics platforms, high-speed networking acquired with Mellanox, and the CUDA software stack that binds the ecosystem together. Headquartered in Santa Clara, California, the company sells to cloud providers, enterprises, research institutions and gamers worldwide and is one of the most valuable listed businesses on the Nasdaq.

NVIDIA is seeking an NCX Engineer, AI Accelerator to join our AI Accelerator team, collaborating closely with strategic customers to implement and enhance groundbreaking AI workloads! You will deliver hands-on technical assistance for advanced AI deployments, intricate distributed systems, and ensure customers realize efficient performance from NVIDIA's AI platform across varied environments. We partner with the world's most innovative AI companies to address their most challenging technical problems.

What you will be doing:

In this role, you will develop innovative solutions that advance AI infrastructure capabilities. You will directly influence customer success with breakthrough AI initiatives.

  • Build and deploy custom AI solutions on NCP and Neo Cloud platforms, including distributed training, inference optimization, and MLOps pipelines constructed on NVIDIA reference architectures.

  • Act as the main technical contact for strategic NCPs, offer remote and on-site support, troubleshoot complex production problems, and guide partner engineering teams on NVIDIA platform guidelines.

  • Deploy and manage AI workloads across DGX Cloud, NCP data centers, and major CSP environments using Kubernetes, containers, and GPU scheduling systems aligned to NCP builds.

  • Profile and tune large-scale training and inference workloads on NCP platforms. Implement observability and SLO/SLA monitoring. Lead detailed efforts to reduce latency, cost, and operational risk.

  • Implement and expand NVIDIA reference architectures on partner platforms, develop integrations with partner control planes and customer environments, and ensure smooth API, data pipeline, and enterprise software connectivity.

  • Build detailed implementation guides, runbooks, and post-mortem documentation that codify standard methodologies for running NVIDIA AI workloads at scale on NCP platforms.

What we need to see:

  • BS, MS, or Ph.D. in Computer Science, Computer/Electrical Engineering, or a related technical field, or equivalent experience.

  • 8+ years of experience in customer facing technical roles such as Solutions Engineering, DevOps, Site Reliability, or ML Infrastructure Engineering, ideally supporting large-scale cloud or service provider environments.

  • Strong expertise in Linux systems, distributed computing, Kubernetes, containers, and GPU scheduling on multi-tenant or service-provider platforms.

  • Demonstrated AI/ML experience supporting large-scale training and inference workloads (e.g., LLMs, generative models, recommendation systems) in production or critically important environments.

  • Solid programming skills in Python/Go, with hands-on experience using frameworks such as PyTorch or TensorFlow for training and serving.

  • Demonstrated capability to collaborate with customer and partner engineering teams in fast-paced environments, guide intricate technical investigations, and bring issues to root cause and resolution.

  • Excellent communication and technical presentation skills, with the ability to clearly articulate architectures, trade-offs, and recommendations to both engineering and leadership audiences.

Ways to stand out from the crowd:

  • Experience with the NVIDIA ecosystem, including DGX systems, CUDA, NeMo, Triton, NIM, and NVIDIA networking technologies such as InfiniBand and RoCE.

  • Direct experience collaborating with NVIDIA Cloud Partners, hyperscale CSPs, or managed AI cloud platforms, including implementation of NVIDIA reference architectures for AI infrastructure.

  • Deep familiarity with MLOps and cloud-native practices: containerization, CI/CD pipelines, observability stacks (Prometheus, Grafana, OpenTelemetry), and GitOps workflows.

  • Background in infrastructure as code (Terraform, Ansible, or similar) for repeatable deployment and configuration of GPU-accelerated clusters and NCP building blocks.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
397,589 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Shanghai
$16k – $42k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
PostgreSQL
Redis
DevOps
Alertmanager
Ansible
Bitbucket
CI/CD
Docker
Git
GitLab
GitLab CI
Grafana
HAProxy
Jenkins
KVM
Loki
Nginx
OpenTofu
Prometheus
TeamCity
Terraform
Management
Confluence
Jira
Apply
$103k – $215k per year (Estimated) • Remote • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
Edge AI
ONNX
PyTorch
Quantization
ROCm
TensorRT
Triton
vLLM
KV Cache
ONNX Runtime
Speculative Decoding
Apply
$82k – $196k per year (Estimated) • Remote • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
Edge AI
ONNX
PyTorch
Quantization
ROCm
TensorRT
Triton
vLLM
KV Cache
ONNX Runtime
Speculative Decoding
Apply
$67k – $160k per year (Estimated) • Remote • Full-Time
Python
AI/ML
CUDA
CUDA Toolkit
Edge AI
ONNX
PyTorch
Quantization
ROCm
TensorRT
Triton
vLLM
KV Cache
ONNX Runtime
Speculative Decoding
Apply
$25k – $67k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Pune
JavaScript
TypeScript
Databases
PostgreSQL
Frontend
Angular
DevOps
CI/CD
Git
Kubernetes
Twelve-Factor App
Apply
In office • Full-Time • Bachelor's Degree • Yokneam
DevOps
HPC
Apply
Release Manager 3 days ago
$105k – $241k per year (Estimated) • In office • Full-Time • 3+ years exp • Master's Degree • Tel Aviv
DevOps
GitHub
Analytics
Power BI
Management
Confluence
Jira
Apply
In office • Full-Time • 5+ years exp • Beijing • Shanghai • Shenzhen
AI/ML
CUDA
CUDA Toolkit
DevOps
HPC
Chips/EDA
PoC Library
Apply
$65k – $228k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Yokneam
Python
Apply
$114k – $274k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • Yokneam • Tel Aviv
Apply
In office • 3+ years exp • Bachelor's Degree • Shanghai
Apply
In office • Full-Time • 10+ years exp • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
ABAP
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Shanghai • Guangzhou • Shenzhen • Dalian • Chengdu
Apply
In office • Internship • 8+ years exp • Shanghai
Apply
In office • Internship • Master's Degree • Beijing • Shanghai • Shenzhen
C++
AI/ML
CUDA
CUDA Toolkit
Speech Recognition
Apply
See all jobs
This is one of many
397,589 more open roles from verified company boards, updated every day.