1,389,218open jobs
80,412companies
207,731added this week
Browse all
Salary
≈ $162k – $358k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Sep 4, 2026. Π scores B on the Alion truth index.

Overview
Company
Impact
Profile match

Π

Join our ML Infrastructure team as a Machine Learning Infrastructure Engineer. You will be responsible for designing, implementing, and maintaining systems for large-scale model training, optimizing performance, and enabling rapid iteration. You will work closely with researchers to scale JAX-based training across TPU and GPU clusters and contribute to core training code. Strong software engineering fundamentals and experience in ML training infrastructure are required.

Missions

  • Conception, mise en œuvre et maintenance des systèmes pour l'entraînement de modèles à grande échelle, y compris la planification, la gestion des tâches, le point de contrôle et la journalisation.
  • Collaboration avec les chercheurs pour étendre l'entraînement basé sur JAX sur des clusters TPU et GPU, en minimisant les frictions.
  • Profilage et amélioration de l'utilisation de la mémoire, de l'utilisation des appareils, du débit et de la synchronisation distribuée.

Profil recherché

- Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms

- Strong cross-functional communication and ownership mindset

- Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)

- Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines

- Ability to debug and optimize performance bottlenecks across the training stack

- Hands-on large-scale training experience in JAX (preferred), PyTorch

- Experience designing abstractions that balance researcher flexibility with system reliability

- Background in robotics, multimodal models, or large-scale foundation models

- Experience operating close to hardware (GPU/TPU performance tuning)

- Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels)

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,389,218 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
AI Engineer Intern 1 day ago
≈ $82k – $245k per year (Estimated) • In office • Internship • Bachelor's Degree • Irving
Python
JavaScript
TypeScript
SQL
Node JS
Databases
PostgreSQL
pgvector
OpenSearch
AI/ML
LangGraph
LangChain
Model Context Protocol
Vertex AI
Embeddings
Function Calling
AI Agents
Langfuse
LangSmith
AWS Bedrock
Gemini
LLM
RAG
OpenAI
Anthropic
LLM Evaluation
LLM Guardrails
Tool Use
Vertex AI Agent Builder
Machine Learning
DevOps
Rest API
GCP
OpenTelemetry
CI/CD
Git
AWS
Docker
GitHub
Apply
≈ $122k – $270k per year (Estimated) • Hybrid • Manchester
Python
SQL
AI/ML
LangChain
Hadoop
Spark
Scikit-learn
TensorFlow
PyTorch
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Management
Agile
Apply
$180k – $230k per year • In office • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
Fine-tuning
Computer Vision
Diffusers
PyTorch
Hugging Face
Post-training
Machine Learning
Apply
$170k – $230k per year • Remote (United States) • Full-Time • 5+ years exp
Python
AI/ML
Claude Code
AI Agents
Accelerate
PyTorch
OpenAI Codex
DevOps
GitHub Actions
CI/CD
Docker
Apply
$150k – $200k per year • In office • Full-Time • 3+ years exp • San Francisco
Python
DevOps
CI/CD
Apply
≈ $42k – $108k per year (Estimated) • In office • Full-Time • Prague
Go
JavaScript
TypeScript
Node JS
Databases
Redis
RabbitMQ
DevOps
Terraform
GCP
CI/CD
AWS
Kubernetes
Platform Engineering
FinOps
GitHub
Apply
≈ $46k – $134k per year (Estimated) • Hybrid • 7+ years exp • Bachelor's Degree • Madrid
Databases
SAP HANA
DevOps
GCP
Azure
AWS
Amazon ECS
Linux
Unix
Apply
≈ $72k – $177k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
SQL
DevOps
GCP
Azure
AWS
FinOps
Analytics
Power BI
Microsoft Excel
Apply
≈ $90k – $198k per year (Estimated) • In office • 8+ years exp • Barcelona
AI/ML
Copilot
LLM
EU AI Act
NIST AI RMF
DevOps
AWS
Windows
Apply
≈ $19k – $49k per year (Estimated) • In office • 2+ years exp • Moscow
Python
C
Python
FastAPI
Aiohttp
C
FFmpeg
AI/ML
Qwen
OpenCV
YOLO
AI Agents
VLM
PyTorch
ResNet
EfficientNet
Triton
CVAT
DevOps
Rest API
gRPC
Docker
Robotics
GStreamer
Apply
≈ $170k – $375k per year (Estimated) • In office • Full-Time • San Francisco
Databases
ClickHouse
AI/ML
Flink
Ray
Machine Learning
Apply
≈ $154k – $322k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
C++
DevOps
WebRTC
eBPF
Robotics
Teleoperation
Apply
≈ $101k – $206k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
Apply
≈ $145k – $291k per year (Estimated) • In office • Full-Time • 2+ years exp • San Francisco
Robotics
Teleoperation
Apply
≈ $93k – $203k per year (Estimated) • In office • Full-Time • 2+ years exp • Associate's Degree • San Francisco
Apply
≈ $164k – $363k per year (Estimated) • In office • Full-Time • San Francisco
Python
C++
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
Quantization
Multimodal AI
SGLang
TensorRT
PyTorch
CUDA
Triton
CUTLASS
Apply
AI Engineer (US) 12 hours ago
$110k – $185k per year • In office • Full-Time • 1+ year exp • San Francisco
Python
AI/ML
LangChain
NLP
PyTorch
LLM
RAG
Hugging Face
Machine Learning
DevOps
Rest API
GCP
Azure
AWS
Apply
$4k – $14k per year • Equity • In office • 2+ years exp • PhD • San Francisco
Python
Go
Java
Kotlin
C++
AI/ML
Machine Learning
DevOps
Kubernetes
Apply
$4k – $14k per year • Equity • In office • 2+ years exp • Bachelor's Degree • San Francisco
Python
Databases
Databricks
AI/ML
Spark
Airflow
Dagster
MLFlow
Transformers
TensorFlow
Pandas
PyTorch
BERT
Machine Learning
DevOps
AWS
Apply
$4k – $14k per year • Equity • In office • 3+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
Cursor
Qwen
DeepSeek
Claude Code
LoRA
Model Context Protocol
vLLM
Fine-tuning
Embeddings
RLHF
Reinforcement Learning
Quantization
AI Agents
SGLang
AWQ
GPTQ
TensorRT
PEFT
TensorRT-LLM
LLM
RAG
OpenAI Codex
DPO
SFT
Post-training
LLM Guardrails
KV Cache
Machine Learning
DevOps
GCP
AWS
Kubernetes
Apply
See all jobs
This is one of many
1,389,218 more open roles from verified company boards, updated every day.