579,193open jobs
24,952companies
79,803added this week
Browse all
Salary
$160k – $200k per year
Location
Remote/Hybrid (Santa Clara, United States)
Seniority
Senior · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
PlusAI is transforming transportation with Physical AI, developing SuperDrive™ virtual driver software for factory-built autonomous commercial vehicles and HyperFoundry™ AI data platform.

As a Senior ML Infrastructure Engineer at Plus, you will design scalable architectures capable of handling petabytes of data while ensuring optimal performance for both training and inference phases. You will build robust pipelines for managing model versioning systems and experiment tracking frameworks, which are essential for maintaining reproducibility across experiments. Additionally, you will be responsible for managing large-scale GPU clusters. This role offers unparalleled opportunities-both technically and professionally-for individuals passionate about solving challenging problems using modern cloud-native technologies. Ideal candidates thrive in environments that leverage tools such as Docker containers orchestrated via Kubernetes clusters, seamlessly integrated with state-of-the-art deep learning frameworks like PyTorch or TensorFlow. If you are eager to push the boundaries of what's possible in machine learning infrastructure and contribute to cutting-edge solutions, this position is an excellent fit!

Responsibilities:

  • Design and develop scalable, high-performance systems for training, inference, deploying, and monitoring ML models at scale.
  • Build and maintain efficient data pipelines, model versioning systems, and experiment tracking frameworks.
  • Collaborate with cross-functional teams, including ML researchers and engineers, to identify bottlenecks and improve platform usability.
  • Implement distributed systems and storage solutions optimized for machine learning workloadsDrive improvements in CI/CD workflows for ML models and infrastructure.
  • Ensure high availability and reliability of the ML platform by implementing robust monitoring, logging, and alerting systems.
  • Stay current with industry trends and integrate relevant tools and frameworks to enhance the platform.
  • Mentor junior engineers and contribute to a culture of technical excellence
  • Ensure that your work is performed in accordance with the company’s Quality Management System (QMS) requirements and contribute to continuous improvement efforts.
  • Ensure team compliance with QMS, monitor quality, and drive process improvements.

Required Skills:

  • Phd or MS in Computer Science, Electrical Engineering, or related field
  • Good oral and written communication skills
  • Phd new grad or Masters with 3+ years of software engineering experience with a focus on ML infrastructure or distributed systems.
  • Proficiency in in Python, C++, SQL
  • Deep understanding of containerization, orchestration technologies, distributed ML workload, and experiment tracking tools (e.g., Docker, Kubernetes, multiprocessing, Kubeflow, and mlflow)
  • Deploy and manage resources across multiple cloud platforms (AWS, GCP, or on-prem environments)
  • Proficiency in at least one deep learning framework, such as PyTorch and data pipeline tools (e.g., Apache Airflow, Prefect).
  • Strong knowledge of distributed systems, databases, and storage solutions.
  • Extensive software design and development skills.
  • Ability to learn and adapt to new technologies and contribute in a productive environment.

Preferred Skills:

  • Familiarity with fundamental deep learning architectures, such as Convolutional Neural Networks (CNNs) and Transformer models
  • Experience in building large-scale ML datasets, MLOps pipelines, and distributed computing frameworks like Ray
  • Experience working with autonomous vehicles or robotics

Salary Range:

  • $160,000 - $200,000 a year
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
579,193 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$101k – $169k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Philadelphia • Saint Louis
Python
SQL
Analytics
Power BI
Alteryx
Microsoft Excel
Apply
$154k – $322k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Boise
Python
JavaScript
TypeScript
C#
C++
Perl
C#
.NET
AI/ML
Model Context Protocol
Prompt Engineering
AI Agents
LLM Guardrails
Frontend
Angular
DevOps
CI/CD
Apply
Business Analyst 8 min ago
$31k – $71k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Monterrey
SQL
Apply
$59k – $138k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Santiago
SQL
Apply
$29k – $83k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Monterrey
SQL
Apply
$110k – $140k per year • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
DevOps
GitHub Actions
CI/CD
Jenkins
Grafana
Robotics
CARLA
Lanelet2
Motion Planning
Apply
$130k – $220k per year • Remote/Hybrid • Full-Time • 3+ years exp • PhD • Santa Clara
Python
C++
Robotics
ROS
Path Planning
Motion Planning
Imitation Learning
Apply
$130k – $220k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Santa Clara
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
TensorRT
PyTorch
CUDA
ONNX Runtime
DevOps
CI/CD
Apply
$130k – $220k per year • Remote/Hybrid • Full-Time • Master's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
Reinforcement Learning
PyTorch
Robotics
Sim-to-Real
Imitation Learning
Reinforcement Learning
Apply
$130k – $220k per year • Remote/Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
Diffusion Models
TensorRT
PyTorch
CUDA
ONNX Runtime
Interpretability
Apply
$46k – $71k per year (Estimated) • In office • Internship • PhD • Santa Clara • Irvine • Austin • Morrisville • Chandler
Python
AI/ML
Copilot
LangGraph
AutoGen
LangChain
Claude
ChatGPT
LlamaIndex
Model Context Protocol
Fine-tuning
Reinforcement Learning
Prompt Engineering
Multimodal AI
Diffusion Models
AI Agents
TensorFlow
PyTorch
CrewAI
LLM
RAG
Hugging Face
A2A
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
Git
Management
n8n
Apply
$61k – $100k per year • In office • Bachelor's Degree • Santa Clara
Apply
$167k – $291k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
Bash
Databases
Azure Cosmos DB
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
Terraform
GCP
CloudFormation
Azure
CI/CD
GitOps
AWS
Kubernetes
Platform Engineering
Service Mesh
Self-Healing
Amazon EKS
Google GKE
Azure AKS
AWS Lambda
Amazon EC2
AIOps
Incident Management
SLI/SLO/SLA
Management
ServiceNow
Apply
$136k – $213k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Apply
$168k – $270k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Python
SQL
AI/ML
Spark
DevOps
Terraform
Azure
AWS
Docker
Kubernetes
Analytics
Tableau
Power BI
ETL/ELT
SAP BusinessObjects
Management
Outlook
Apply
See all jobs
This is one of many
579,193 more open roles from verified company boards, updated every day.