368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$350k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design and build the core systems that enable efficient large-scale model training with a focus on numerics. You will focus on improving the numerical foundations of our distributed training stack, from precision formats and kernel optimizations to communication frameworks that make training trillion-parameter models stable, scalable, and fast.

This role is ideal for someone who thrives at the intersection of research and systems engineering: a builder who understands both the math of optimization and the realities of distributed compute.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Design and optimize distributed training infrastructure for large-scale LLMs, focusing on performance, stability, and reproducibility across multi-GPU and multi-node setups.

  • Implement and evaluate low-precision numerics (for example, BF16, MXFP8, NVFP4) to improve efficiency without sacrificing model quality.

  • Develop kernels and communication primitives that use hardware-level support for mixed and low-precision arithmetic.

  • Collaborate with research teams to co-design model architectures and training recipes that align with emerging numeric formats and stability constraints.

  • Prototype and benchmark scaling strategies such as data, tensor, and pipeline parallelism that integrate precision-adaptive computation and quantized communication.

  • Contribute to the design of our internal orchestration and monitoring systems to ensure that thousands of distributed experiments can run efficiently and reproducibly.

  • Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases in areas such as floating-point numerics, low-precision arithmetic, and distributed systems.

Preferred qualifications - we encourage you to apply if you meet some but not all of these:

  • Familiarity with distributed frameworks such as PyTorch/XLA, DeepSpeed, Megatron-LM.

  • Experience implementing FP8, INT8, or block-floating point (MX) formats and understanding their numerical trade-offs.

  • Prior contributions to open-source deep learning infrastructure such as PyTorch, DeepSpeed, or XLA.

  • Publications, patents, or projects related to numerical optimization, communication-efficient training, or systems for large models.

  • Experience training and supporting large-scale AI models.

  • Track record of improving research productivity through infrastructure design or process improvements.

Logistics

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$105k – $175k per year • In office • Full-Time • Raleigh
Python
AI/ML
AI Agents
Amazon SageMaker
AutoGen
AWS Bedrock
CrewAI
Edge AI
Function Calling
LangChain
LangGraph
Model Context Protocol
NLP
Prompt Engineering
PyTorch
RAG
TensorFlow
DevOps
Amazon EC2
Amazon ECS
Amazon EKS
Amazon S3
AWS
AWS Lambda
Azure
GCP
Git
Kubernetes
Apply
$14k – $33k per year (Estimated) • Remote • Full-Time • Moscow
SQL
Java
Java
Spring Cloud
Databases
Apache Kafka
PostgreSQL
AI/ML
PyTorch
TensorFlow
DevOps
API Gateway
Docker
Git
Grafana
Jenkins
Kubernetes
Prometheus
Rest API
QA
Insomnia
Pact
Postman
Selenium
SoapUI
Apply
$175k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Master's Degree • New York
C++
Python
SQL
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
Scikit-learn
TensorFlow
Apply
DS/ AI Lead 7 hours ago
$49k – $118k per year (Estimated) • In office • Full-Time • Moscow
Python
AI/ML
AutoGen
Chain-of-Thought
CrewAI
DeepSpeed
Fine-tuning
GigaChat
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Transformers
Apply
AI / ML Engineer 8 hours ago
$45k – $98k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad
AI/ML
PyTorch
TensorFlow
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco
Python
Rust
TypeScript
JavaScript
AI/ML
Fine-tuning
Frontend
React.js
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
Fine-tuning
LoRA
PEFT
DevOps
CI/CD
Kubernetes
SRE
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
CUDA
CUDA Toolkit
NCCL
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.