703,353open jobs
41,539companies
100,398added this week
Browse all
Salary
$350k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

Mid-training is a step between pre-training and post-training, where we take a base model and train it into the foundation for reasoning. This role owns the late-stage training responsibility that shape what our models are fundamentally capable of, including things like synthetic data strategies, the data mix, quality uplift, context extension and capabilities across coding, math, reasoning, and so on.

This role blends fundamental research and practical engineering, as we do not distinguish between the two internally. It's an excellent fit for someone comfortable working across the boundary of pre-training and post-training, and who wants to shape what our models can do at their core.

What You’ll Do

  • Own the data. Decide what the model needs to see for each capability and each area of knowledge, then source, curate, and synthesize it. Build the pipelines that filter, deduplicate, verify, and rewrite raw material into training-grade datasets. You'll be responsible for producing data that actually moves the model, and for making sure it reflects how people really use these models, not just clean benchmark-style tasks.

  • Improve what the model knows. Design and measure interventions that increase knowledge: targeted corpora, synthetic rephrasings and QA over source documents, knowledge-dense mixes. Characterize how knowledge scales with data and when it is retained through post-training versus forgotten.

  • Instill behaviors and set the prior. Introduce new behaviors during mid-training and measure them the right way: not just whether an eval goes up, but whether post-training becomes easier, more sample-efficient, and more stable as a result. Work closely with the post-training team to decide which behaviors belong in mid-training and which belong in RL.

  • Build the quality pipeline. Own the automatic filtering and scoring stack: train and calibrate quality classifiers and LLM-based judges, build verifiers for synthetic data, and use them to raise the quality bar of the mix by a lot, not a little. Give the team fine-grained control over data attributes (difficulty, domain, format, style, correctness) and measure the effect of each.

  • Develop and tune the recipe. Iterate on mid-training recipes: the collection of datasets, training stages, annealing schedules, and hyperparameters. Measure how recipe choices affect metrics, including downstream of post-training.

  • Iterate on evals. Mid-training involves a never-ending loop of defining a set of evaluations, optimizing them, and then realizing your existing evals don't capture what matters. You'll be responsible for both making numbers go up and making sure the numbers are meaningful.

  • Debug and understand. While tuning the details of a training configuration, we often observe results that don't quite make sense. You'll be responsible for both getting things to work and developing a deeper understanding we can bring to the next problem.

  • Scale and explore. Mid-training will involve a combination of scaling existing methodologies and developing new ones. We'll want to both measure how performance scales with dataset size and explore using completely different kinds of training data.

Skills and Qualifications

Required qualifications:

  • Proficiency in Python and familiarity with deep learning frameworks (e.g., PyTorch, TensorFlow, or JAX). Comfort debugging distributed training and writing code that scales.

  • Bachelor’s degree or equivalent experience in Computer Science, Machine Learning, Physics, Mathematics, or a related discipline with strong theoretical and empirical grounding.

  • Clarity in communication, an ability to explain complex technical concepts in writing.

Preferred qualifications (we encourage you to apply if you meet some but not all of these):

  • A strong grasp of probability, statistics, and ML fundamentals. You can look at experimental data and distinguish between real effects, noise, and bugs.

  • Experience building or owning training datasets for large models, including both synthetic data pipelines and real user data (collection, curation, filtering, or mixture design).

  • Experience with knowledge injection, continued pre-training, or domain adaptation of large models, and with measuring what a model knows.

  • Experience building model-based data quality systems: quality classifiers, LLM judges, or verification and rewriting pipelines operating at the scale of billions of tokens.

  • Experience generating synthetic data and understanding its failure modes (diversity collapse, contamination, teacher artifacts).

  • Research or engineering contributions in alignment, data-centric AI, or human-AI collaboration.

Logistics

  • This is an IC role within Thinking Machines’ Research Team. You will report directly to our Head of Research, and work most closely with our small group of post-training researchers.

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
703,353 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$120k per year • In office • Internship • PhD • Mason
Python
AI/ML
LangGraph
LangChain
AI Agents
Arize Phoenix
TensorFlow
PyTorch
LLM
Time Series Forecasting
Agentic Workflows
Multi-Agent Systems
Machine Learning
Apply
$179k – $205k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • McLean • Richmond • Chicago • Plano • San Francisco
Python
JavaScript
TypeScript
SQL
Node JS
Scala
AI/ML
LangGraph
LangChain
AI Agents
LLM
RAG
Agentic Workflows
Multi-Agent Systems
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Apply
$53k – $100k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Milwaukee
Python
DevOps
CI/CD
TCP/IP
Cybersecurity
Wireshark
Nmap
Management
Agile
Apply
$73k per year • Equity • In office • Full-Time • Goleta
Python
C#
SPSS
Python
Hypothesis
Design
SolidWorks
Management
Outlook
Microsoft Office
Apply
$157k – $227k per year • In office • Full-Time • 7+ years exp • Master's Degree • Wilmington
Python
Verilog
C++
VHDL
MATLAB
MATLAB
Simulink
DevOps
Linux
Chips/EDA
HDL Coder
Apply
$300k – $360k per year • In office • Full-Time • 10+ years exp • San Francisco • Washington
AI/ML
Fine-tuning
EU AI Act
Apply
Product Policy 3 days ago
$300k – $360k per year • In office • Full-Time • 10+ years exp • San Francisco
AI/ML
Fine-tuning
AI Agents
Recommender Systems
Apply
$350k – $475k per year • Remote/Hybrid • Full-Time • PhD • San Francisco
Python
AI/ML
LoRA
Fine-tuning
RLHF
JAX
PEFT
TensorFlow
PyTorch
Post-training
RLAIF
Machine Learning
Apply
$475k – $530k per year • Remote/Hybrid • Full-Time • PhD • San Francisco
Python
AI/ML
LoRA
Fine-tuning
RLHF
JAX
PEFT
TensorFlow
PyTorch
Post-training
RLAIF
Machine Learning
Apply
$350k – $475k per year • Remote/Hybrid • Full-Time • PhD • San Francisco
Python
Clarity
AI/ML
vLLM
Fine-tuning
Quantization
JAX
SGLang
TensorFlow
PyTorch
LLM
Post-training
Machine Learning
Apply
$135k – $245k per year • In office • Full-Time • 7+ years exp • PhD • San Francisco • Chicago • New York • Boston • Washington
AI/ML
AI Agents
Agentforce
Agentic Workflows
Management
ServiceNow
Apply
$197k – $314k per year • In office • Full-Time • 10+ years exp • PhD • San Francisco • Boston • Chicago • New York • Austin
AI/ML
AI Agents
LLM
Agentforce
LLM Guardrails
Apply
$130k – $217k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
Apply
$90k per year • In office • Full-Time • Bachelor's Degree • San Francisco • New York • Dallas
Apply
$100k – $186k per year • Equity • Remote • Full-Time • 5+ years exp • New York • Ann Arbor • San Francisco • Frisco • Chicago
Apply
See all jobs
This is one of many
703,353 more open roles from verified company boards, updated every day.