664,807open jobs
38,858companies
99,669added this week
Browse all
Salary
$162k – $310k per year (Estimated)
Location
In office (San Jose)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Skip to the content About Contact Home · Product Design · Education · Online Class · Case Study: BIDS Trading · Research · Development Home About Contact Twitter LinkedIn Email Interface Innovation since 1983 We provide rigorous design services to the financial community, the re/insurance industry and other knowledge-work fields.

About the Company

DiDi's autonomous driving unit was established in 2016 with the mission of developing Level 4 autonomous driving (AD) technology to make transportation safer and more efficient. In August 2019, the unit became an independent company, DiDi Autonomous Driving, dedicated to advanced AD R&D, product application, and business expansion. We believe integrating AD technology into a shared-mobility fleet will generate immense social value. By leveraging DiDi's specialized technology, operational expertise, and integrated ecosystem, we are positioned to build and operate a highly efficient, user-oriented autonomous fleet.

About The Role

We are seeking an experienced and mission-driven Senior/Sr. Staff AI Infrastructure Engineer, Inference & Optimization to lead the performance tuning, deployment, and resource scheduling of cutting-edge AI models across on-vehicle and cloud infrastructure. In this role, you will design high-efficiency inference pipelines, build system-level stability frameworks, and optimize hardware execution to ensure ultra-low latency and rock-solid operational reliability. You will act as a technical leader in AI infrastructure, accelerating model iteration and bridging the gap between frontier deep learning algorithms and real-time autonomous systems.

Responsibilities

  • Own the deployment, optimization, and resource scheduling of vehicle-side AI models, ensuring high efficiency, low latency, and robust execution within embedded constraints.

  • Lead vehicle-side system stability initiatives, conducting independent root-cause analysis and driving resolution for complex, system-level performance bottlenecks and runtime anomalies.

  • Architect and scale service-oriented deployment environments for Large Language Models (LLMs) and foundational models to support offline simulation, automated annotation, and rapid model validation.

  • Track and evaluate cutting-edge industry methodologies, continuously integrating advanced optimization toolchains, quantization techniques, and execution engines.

  • Establish system-level profiling and telemetry frameworks using CUDA tools to monitor, analyze, and maximize hardware utilization across target GPU architectures.

  • Collaborate cross-functionally with Autonomous Driving Perception/Prediction, Cloud Infrastructure, and Safety teams to enable rapid algorithm iteration and scalable vehicle deployment.

Qualifications

  • Master’s or higher degree in Computer Science, Software Engineering, Systems Engineering, or a closely related technical field.

  • 3-8+ years of industry experience in high-performance computing, AI infrastructure, model optimization, or embedded deployment.

  • Strong proficiency in C++ and Python, with solid expertise in parallel programming (CUDA, OpenMP) and low-level system profiling tools.

  • Deep familiarity with mainstream inference engines (e.g., TensorRT, ONNX Runtime) and specialized LLM inference/serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM).

  • Practical understanding of modern GPU hardware architectures (e.g., NVIDIA Hopper, Thor) and memory bandwidth management.

  • Demonstrated ability to diagnose complex software-hardware integration issues and drive scalable, production-grade solutions.

Preferred Qualifications

  • Hands-on experience optimizing and deploying AI models on the NVIDIA Thor platform, including hardware resource scheduling and acceleration.

  • Proven track record of serving large foundation models (e.g., LLaMA, Qwen, GPT) in production or high-throughput cloud pipelines using frameworks like vLLM, SGLang, TGI, or LightLLM.

  • Background in deep learning training frameworks (PyTorch) and practical experience with model quantization (INT8/FP8/AWQ), kernel fusion, or graph compilation.

  • Experience deploying real-time, high-availability AI workloads in autonomous vehicles, robotics, or edge devices.

The base salary range for this full-time position is $169,783 - $351,000 annually in addition to bonus, equity and benefits. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.

I acknowledge that prior to submitting this application, I have read and accepted the Privacy Notice for California Residents which is available on https://v.didi.cn/AQnxlBa

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
664,807 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$22k – $54k per year (Estimated) • Remote • Full-Time • Moscow
Python
C++
Bash
Python
bandit
C++
CMake
DevOps
GitLab CI
CI/CD
Jenkins
Cybersecurity
SonarQube
Semgrep
SBOM
Apply
up to $22k per year (net) • In office • Full-Time • Saint Petersburg
Python
SQL
Python
SQLAlchemy
FastAPI
Ruff
AI/ML
Claude
ChatGPT
Model Context Protocol
AI Agents
LLM
Apply
$19k – $38k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Moscow
Python
SQL
Python
pySpark
Databases
MS SQL
Greenplum
AI/ML
Hadoop
Spark
GigaChat
Analytics
QlikSense
Apply
Data Engineer 2 days ago
$21k – $57k per year (Estimated) • In office • 1+ year exp • Bengaluru
Python
SQL
Databases
Snowflake
Apache Kafka
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Spark
dbt
DevOps
CI/CD
AWS
Analytics
ETL/ELT
AWS Glue
Apply
$17k – $40k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Mumbai • Bengaluru
Python
SQL
Analytics
Tableau
Matplotlib
Looker
Apply
$43k – $86k per year (Estimated) • In office • Internship • PhD • San Jose
C++
AI/ML
AI Agents
Robotics
Motion Planning
Apply
$149k – $247k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Jose
Apply
$187k – $310k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
C++
C++
CMake
LLVM
DevOps
eBPF
Apply
$129k – $271k per year (Estimated) • In office • Full-Time • Bachelor's Degree • San Jose
Python
C++
AI/ML
World Models
Robotics
Motion Planning
Apply
$147k – $280k per year (Estimated) • In office • Full-Time • 10+ years exp • San Jose
Python
JavaScript
C++
C++
PyTorch C++
AI/ML
Copilot
DeepSpeed
vLLM
Triton Inference Server
AI Agents
Llama
PyTorch
LLM
Ray
Hugging Face
LLMOps
Megatron-LM
Multi-Agent Systems
DevOps
Kubernetes
Apply
Hardware Engineer 1 day ago
$144k – $230k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Jose
Apply
ASIC DFT Engineer 1 day ago
$110k – $176k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Fort Collins • San Jose
Verilog
DevOps
Vector
Chips/EDA
Tessent
Apply
$267k – $314k per year • Remote/Hybrid • Full-Time • 20+ years exp • Bachelor's Degree • San Jose
Management
Agile
Apply
$162k – $315k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Jose • Folsom • Boise
Python
AI/ML
Model Context Protocol
XGBoost
Scikit-learn
AI Agents
TensorFlow
PyTorch
RAG
Knowledge Graph
LLM Guardrails
Agentic Workflows
DevOps
GCP
OpenShift
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$119k – $172k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • San Jose
Marketing
Salesforce
Apply
See all jobs
This is one of many
664,807 more open roles from verified company boards, updated every day.