782,981open jobs
49,348companies
120,575added this week
Browse all
Salary
≈ $160k – $309k per year (Estimated)
Location
In office (San Francisco)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on May 18, 2026. Maker Maker AI scores B on the Alion truth index.

Overview
Company
Impact
Profile match

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You build and operate the inference systems that serve our models in production. The work spans serving infrastructure, runtime optimization, and the long tail of production infrastructure that come with running real workloads.

This is an engineering role, not a research role. You'll measure, profile, debug, and ship. You'll work alongside researchers, but your job is to make their work fast and reliable in production. Real ownership, real autonomy.

WHAT YOU'LL DO

  • Build, operate, and harden production inference systems serving large models at high throughput

  • Own the performance characteristics of those systems end-to-end: throughput, latency, cost-per-token, reliability under load

  • Profile real workloads to identify bottlenecks; ship fixes that move the metric you set out to improve

  • Implement and integrate inference optimizations from the research team (quantization, custom kernels, scheduling improvements, memory management) into production

  • Design observability into the inference layer: metrics, tracing, alerting that surface regressions before users notice them

  • Run capacity planning, autoscaling, and load testing for varied workload shapes (batch, online, mixed, agentic)

  • Diagnose and resolve production incidents; write postmortems that turn bugs into systemic fixes

WHAT WE'RE LOOKING FOR

  • Senior ML systems engineer with 3+ years building production-grade, large-scale serving infrastructure

  • Strong distributed systems experience ; you've been on-call for systems that matter

  • Performance profiling and optimization fluency: you read flame graphs, you are analytical and measured before you change

  • Experience with GPU-accelerated inference at scale (multi-GPU, multi-node, batched and streaming workloads), preferably experience with AMD GPUs

  • Fluent Python; comfortable reading and writing systems-level code in at least one of the following languages: C++,CUDA, ROCm or Triton

  • Track record of shipping production infrastructure, preferably surfaces serving millions of requests across diverse workloads

  • Good written communication; you can write a runbook that someone else can follow at 3am

NICE TO HAVE

  • Open-source contributions to inference / serving frameworks

  • Experience with mixed cloud and on-premises deployments

  • Familiarity with hardware-aware optimization (memory hierarchy, NCCL/RDMA, NUMA)

  • Background in compilers, runtimes, or accelerator software stacks

  • THIS ROLE IS PROBABLY NOT FOR YOU IF

  • You're primarily a researcher, the work here is building, not exploring

  • You want to focus narrowly on one component; this role spans the stack

  • Production responsibility (incidents, on-call, ownership of running systems) isn't appealing

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
782,981 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
≈ $154k – $281k per year (Estimated) • In office • Full-Time • Master's Degree • San Jose
Python
C++
AI/ML
Computer Vision
NeRF
Apply
≈ $149k – $271k per year (Estimated) • In office • Full-Time • PhD • Seattle
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
Databases
MySQL
PostgreSQL
AI/ML
XGBoost
Reinforcement Learning
Scikit-learn
Multimodal AI
LightGBM
TensorFlow
PyTorch
LLM
Machine Learning
DevOps
GCP
Azure
AWS
Apply
≈ $107k – $242k per year (Estimated) • In office • Full-Time • PhD • Seattle
Python
C++
AI/ML
LLM
NCCL
Apply
≈ $108k – $244k per year (Estimated) • In office • Full-Time • PhD • San Jose
Python
C
C++
C
FFmpeg
C++
TensorFlow C++
PyTorch C++
AI/ML
OpenCV
Transformers
TensorFlow
PyTorch
DevOps
Linux
Analytics
A/B Testing
Apply
≈ $148k – $269k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
C++
AI/ML
CUDA Toolkit
CUDA
NCCL
KV Cache
DevOps
eBPF
Linux
Apply
Data Engineer (BI) 1 day ago
≈ $74k – $165k per year (Estimated) • Remote (likely United States) • 3+ years exp • Bachelor's Degree
Python
SQL
Databases
Databricks
MS SQL
AI/ML
Copilot
Spark
DevOps
Azure DevOps
Azure
Platform Engineering
Cybersecurity
HIPAA
Microsoft Entra ID
Analytics
Power BI
Management
Agile
Scrum
Apply
$89k – $121k per year • Remote (United States) • Full-Time • 5+ years exp • PhD • United States
Python
SQL
Databases
Databricks
Cybersecurity
HIPAA
Analytics
Power BI
ETL/ELT
SSIS
Management
Power Automate
Power Apps
Apply
≈ $35k – $88k per year (Estimated) • In office • Full-Time • London
Python
SQL
Analytics
Power BI
Apply
$162k – $190k per year • Hybrid • Full-Time • 10+ years exp • Charlotte
Python
SQL
Databases
Snowflake
AI/ML
dbt
Tokenization
Machine Learning
DevOps
Azure DevOps
GitHub Actions
Prometheus
Azure
CI/CD
Cortex
FinOps
SLI/SLO/SLA
Analytics
ETL/ELT
Fivetran
Azure Data Factory
Data Vault
Apply
$124k – $152k per year • In office • Full-Time • 3+ years exp • PhD • Maplewood
Python
AI/ML
Machine Learning
DevOps
HPC
Apply
≈ $176k – $320k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
Python
AI/ML
JAX
PyTorch
Ray
Multi-Agent Systems
Machine Learning
DevOps
SLURM
Kubernetes
Apply
≈ $176k – $320k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
Python
AI/ML
JAX
PyTorch
Ray
Multi-Agent Systems
Machine Learning
DevOps
SLURM
Kubernetes
Apply
≈ $184k – $333k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
Quantization
Knowledge Distillation
PyTorch
Mixture of Experts
Post-training
Speculative Decoding
Multi-Agent Systems
Model Distillation
Machine Learning
Apply
≈ $176k – $319k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
Function Calling
Chain-of-Thought
AI Agents
SFT
Multi-Agent Systems
Tool Use
Machine Learning
Apply
≈ $189k – $343k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
RLHF
Reinforcement Learning
PyTorch
Synthetic Data
DPO
SFT
Post-training
Multi-Agent Systems
RLAIF
Reward Modeling
Machine Learning
Apply
$125k – $175k per year • Equity 0.2–2% • In office • Full-Time • 3+ years exp • PhD • San Francisco
Python
AI/ML
Multimodal AI
Time Series Forecasting
Machine Learning
Robotics
Sensor Fusion
Apply
$180k – $250k per year • Remote (United States) • Full-Time • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
≈ $134k – $359k per year (Estimated) • In office • Internship • San Francisco
Python
TypeScript
Python
Hypothesis
AI/ML
AI Agents
Post-training
Machine Learning
Apply
Founding Engineer 9 hours ago
$110k – $180k per year • Equity 0.1–1% • In office • Full-Time • San Francisco
Python
JavaScript
Node JS
AI/ML
Vertex AI
OpenAI
Anthropic
Frontend
Next.js
React.js
DevOps
Azure
Kubernetes
Apply
$60k – $84k per year • In office • Internship • San Francisco
Python
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
LLM
OpenAI
Anthropic
LLM Guardrails
Tool Use
Frontend
Next.js
React.js
DevOps
Azure
AWS
Kubernetes
Apply
See all jobs
This is one of many
782,981 more open roles from verified company boards, updated every day.