368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$350k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

Our team scales reinforcement learning for frontier models. Progress in RL is increasingly set by how well it scales: more rollouts, larger models, and training loops that keep large fleets of accelerators doing useful work. We are particularly interested in people working on high-training-compute, long-horizon RL. We believe the biggest gains come from designing the training recipe and the infrastructure together rather than separately, and we are hiring a researcher who wants to own that boundary.

A center of gravity for this role is asynchronous RL. Decoupling generation from training changes both the systems design and the learning problem, and doing it well requires a deep understanding of async RL algorithms, design choices, and trade-offs on both the ML and the systems sides. We expect much of the headroom in RL scaling to come from here.

Because generation dominates the cost of RL at scale, good knowledge of inference systems, low-precision numerics, and quantization is recommended: you should be able to reason quantitatively about rollout throughput and cost (batching, KV cache, MoE serving, speculative decoding) and about how inference constraints shape training design.

This is a research role with full-stack ownership, from the algorithms to the parallelism plan to the health of the run.

What You’ll Do

  • Co-design the RL recipe and the systems that run it: make recipe-level choices jointly with systems-level ones and validate them at frontier scale.

  • Advance asynchronous RL algorithms.

  • Improve the efficiency of rollout generation and its integration with training, treating inference as a first-class part of the RL loop.

  • Run frontier-scale RL end to end: bring up new models and training setups, keep large runs stable and healthy.

  • Jointly optimize the compute and training efficiency of RL: accelerator utilization, memory, communication, and low-precision numerics.

  • Do careful empirical science: ablations and scaling studies backed by instrumentation you can trust, written up clearly.

Skills and Qualifications

Minimum qualifications:

  • Proficiency in Python and familiarity with at least one deep learning framework (e.g., PyTorch, TensorFlow, or JAX). Comfortable with debugging distributed training and writing code that scales.

  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.

  • Clarity in communication, an ability to explain complex technical concepts in writing.

  • Strong research judgment: clean ablations, honest baselines, and clear technical writing.

Preferred qualifications - we encourage you to apply if you meet some but not all of these:

  • PhD in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding; or, equivalent industry research experience.

  • Strong grounding in RL for large language models, such as modern policy optimization methods and their behavior at scale.

  • Deep understanding of asynchronous RL: the algorithms, design choices, and trade-offs, on both the ML and the systems sides.

  • Experience training large models across many accelerators, with comfort inside the distributed stack (parallelism strategies, memory, communication).

  • Good working knowledge of inference systems: able to reason quantitatively about rollout generation throughput and cost.

  • Experience building or operating decoupled generation/training RL systems at scale.

  • Experience with RL on verifiable and agentic tasks, including multi-turn environments.

  • Experience with RL training stability techniques for large runs.

  • Familiarity with low-precision training and inference: numerics, quantization, and their implications for RL.

  • Hands-on work with LLM serving stacks (e.g., SGLang, vLLM, TokenSpeed, or custom engines).

  • Experience with scaling studies for large models.

  • Contributions to open-source training or inference frameworks.

Logistics

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$77k – $190k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Toronto
JavaScript
Python
DevOps
Bitbucket
Git
Apply
$73k – $186k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Toronto
Java
Python
SQL
Databases
Databricks
AI/ML
Spark
DevOps
AWS
Azure
Bitbucket
GCP
Git
Analytics
ETL/ELT
Apply
$88k – $224k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Waterloo
Java
Python
SQL
Databases
Databricks
AI/ML
Spark
DevOps
AWS
Azure
Bitbucket
GCP
Git
Analytics
ETL/ELT
Apply
Senior AI Architect 8 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$55k – $130k per year (Estimated) • Remote • Full-Time • Canada
Java
Python
SQL
Databases
MS SQL
Analytics
Power BI
Tableau
Management
Power Apps
Power Automate
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco
Python
Rust
TypeScript
JavaScript
AI/ML
Fine-tuning
Frontend
React.js
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
Fine-tuning
LoRA
PEFT
DevOps
CI/CD
Kubernetes
SRE
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
CUDA
CUDA Toolkit
NCCL
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 2 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.