825,038open jobs
53,179companies
135,110added this week
Browse all
Salary
≈ $184k – $333k per year (Estimated)
Location
In office (San Francisco)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on May 18, 2026. Maker Maker AI scores B on the Alion truth index.

Overview
Company
Impact
Profile match

ABOUT THE COMPANY

We're building autonomous research agents for recursive self-improvement (multi-agent systems that propose, run, and analyze machine learning experiments). We're a small team based in San Francisco, on-site

ABOUT THE ROLE

You'll be researching making models efficient: quantization, speculative decoding, sparse and structured attention, distillation, mixture-of-experts inference, and the training-time techniques that make those methods possible. The work spans algorithm design, careful evaluation, and pushing methods to where they actually run.

This is a senior research role with a clear engineering edge. You'll spend time at the intersection of model architecture and inference performance, designing methods that move accuracy/latency/cost trade-offs in our favor (then partnering with engineers to make those wins real in production).

WHAT YOU'LL DO

  • Research and develop quantization methods: post-training quantization, quantization-aware training, mixed-precision regimes, low-bit-width arithmetic

  • Design and evaluate speculative decoding approaches: draft models, tree attention, parallel speculation, lookahead decoding

  • Investigate training-time efficiency methods that compose well with inference: distillation, sparse attention, mixture-of-experts, low-rank adaptation, pruning

  • Run controlled experiments at production scale; characterize what works on real workloads, not just toy benchmarks

  • Co-design methods with the inference engineering team: push results to where they actually run, not stop at the paper

  • Read deeply across the efficient ML / efficient inference literature; translate the most useful ideas into our stack

  • Publish when the work warrants it; share findings internally

  • Partner with model and training researchers so efficiency choices align with model architecture and post-training decisions

WHAT WE'RE LOOKING FOR

  • Strong track record of ML research on efficiency methods: quantization, speculative decoding, distillation, MoE, sparse attention, or adjacent

  • 5+ years of hands-on research experience

  • Deep familiarity with both training and inference performance characteristics

  • Fluent in PyTorch, Jax or equivalent; comfortable working at the kernel and serving-framework level when methods require it

  • Track record of moving efficiency research from prototype to production

  • Strong statistical expertise: you'd notice a flawed comparison before someone else points it out

  • Strong written communication

  • Published research at NeurIPS, ICML, ICLR, MLSys, or comparable venues

NICE TO HAVE

  • PhD in ML, systems, or related field

  • Open-source contributions to quantization, speculative-decoding, or efficient-inference libraries

  • Experience with hardware-aware optimization and accelerator-specific tooling

  • Background in numerical methods, low-precision arithmetic, or

  • approximate computation

THIS ROLE IS PROBABLY NOT FOR YOU IF

  • You want to focus on pretraining large models from scratch (that's a different role)

  • You prefer abstract algorithmic research without hands-on implementation

  • You want a fixed benchmark with stable targets (our targets shift with what our models actually need to do)

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
825,038 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$133k – $338k per year • Hybrid • Full-Time • 12+ years exp • Associate's Degree • Dallas • Tampa • Atlanta • Columbus • Houston
Python
Java
AI/ML
LangGraph
AutoGen
LangChain
Claude
Vertex AI
AI Agents
AgentOps
CrewAI
LLM
OpenAI
Anthropic
LLMOps
OpenAI Agents SDK
AWS Bedrock AgentCore
Knowledge Graph
Vertex AI Agent Builder
DevOps
Azure
Apply
$94k – $294k per year • Hybrid • Full-Time • 12+ years exp • Associate's Degree • Dallas • Tampa • Atlanta • Columbus • Houston
Python
Java
Databases
Databricks
Delta Lake
AI/ML
LangGraph
AutoGen
LangChain
Claude
Spark
MLFlow
Vertex AI
AI Agents
CrewAI
LLM
RAG
OpenAI
Anthropic
LLMOps
LLM Evaluation
Agentic Workflows
Multi-Agent Systems
DevOps
CI/CD
Apply
≈ $140k – $245k per year (Estimated) • Remote (United States) • Public Trust • Full-Time • Bachelor's Degree • Salt Lake City
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Microsoft Fabric
AI/ML
Spark
MLFlow
XGBoost
Scikit-learn
PyTorch
RAG
Machine Learning
DevOps
Terraform
Azure
CI/CD
Git
Cybersecurity
HIPAA
FedRAMP
Microsoft Entra ID
Analytics
Power BI
ETL/ELT
SSIS
Azure Data Factory
SSAS
Apply
≈ $145k – $263k per year (Estimated) • In office • Austin
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
$120k – $170k per year • In office • Chicago
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
≈ $49k – $109k per year (Estimated) • In office • 10+ years exp • Bengaluru
Python
AI/ML
LangGraph
LangChain
LoRA
Model Context Protocol
Fine-tuning
Embeddings
Quantization
Function Calling
AI Agents
NLP
PEFT
Transformers
TensorFlow
PyTorch
LLM
RAG
Reranking
Hybrid Search
Anomaly Detection
Hugging Face
GraphRAG
Human-in-the-Loop
Context Engineering
Knowledge Graph
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Kubernetes
Chips/EDA
PoC Library
Apply
≈ $43k – $102k per year (Estimated) • In office • 7+ years exp • Bengaluru
Julia
Julia
SciML
Databases
Databricks
AI/ML
LangGraph
AutoGen
LangChain
Quantization
JAX
AI Agents
ONNX
TensorRT
AWS Bedrock
TensorFlow
PyTorch
CrewAI
RAG
Stable-Baselines3
Ray
RLlib
Triton
PPO
Machine Learning
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Azure AKS
Robotics
MuJoCo
Apply
≈ $84k – $225k per year (Estimated) • In office • Full-Time • United Arab Emirates
Python
JavaScript
TypeScript
C++
AI/ML
Fine-tuning
Reinforcement Learning
Post-training
Machine Learning
Apply
In office • Full-Time • Seongnam
Python
C++
Cython
C++
PyTorch C++
Cython
PyBind11
AI/ML
CUDA Toolkit
Pandas
PyTorch
CUDA
DevOps
Linux
Apply
In office • Full-Time • PhD • Seongnam
Python
Java
C++
Scala
C++
PyTorch C++
Databases
GraphDB
AI/ML
Spark
vLLM
CUDA Toolkit
SGLang
Pandas
PyTorch
LLM
CUDA
DevOps
Prometheus
Cortex
Linux
Apply
≈ $176k – $320k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
Python
AI/ML
JAX
PyTorch
Ray
Multi-Agent Systems
Machine Learning
DevOps
SLURM
Kubernetes
Apply
≈ $176k – $320k per year (Estimated) • In office • Full-Time • 6+ years exp • San Francisco
Python
AI/ML
JAX
PyTorch
Ray
Multi-Agent Systems
Machine Learning
DevOps
SLURM
Kubernetes
Apply
≈ $189k – $343k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
RLHF
Reinforcement Learning
PyTorch
Synthetic Data
DPO
SFT
Post-training
Multi-Agent Systems
RLAIF
Reward Modeling
Machine Learning
Apply
RESEARCHER (GENERAL) 4 months ago
≈ $176k – $320k per year (Estimated) • In office • Full-Time • 5+ years exp • PhD • San Francisco
AI/ML
Function Calling
AI Agents
PyTorch
Multi-Agent Systems
Tool Use
Machine Learning
Apply
INFERENCE ENGINEER 4 months ago
≈ $160k – $309k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Python
C++
AI/ML
CUDA Toolkit
Quantization
AI Agents
CUDA
Triton
NCCL
ROCm
Multi-Agent Systems
Machine Learning
Apply
$130k – $175k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Apply
≈ $105k – $207k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$140k – $295k per year • In office • 5+ years exp • San Francisco
Apply
$100k – $125k per year • In office • 2+ years exp • San Francisco
Python
Apply
$145k – $235k per year • In office • 3+ years exp • San Francisco
Apply
See all jobs
This is one of many
825,038 more open roles from verified company boards, updated every day.