368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$350k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design and build the core systems that enable scalable, efficient training of large models for deployment and research. Your goal is to make experimentation and training at Thinking Machines fast and reliable to ensure our research teams can focus on science, not system bottlenecks.

This role is ideal for someone who blends deep systems and performance expertise with a curiosity for machine learning at scale. You’ll take ownership of the training stack end to end, ensuring every GPU cycle drives scientific progress.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Design, implement, and optimize distributed training systems that scale across thousands of GPUs and nodes for large-scale training workloads.

  • Develop high-performance optimizations to maximize throughput and efficiency.

  • Develop reusable frameworks and libraries to improve training reproducibility, reliability, and scalability for new model architectures.

  • Establish standards for reliability, maintainability, and security, ensuring systems are robust under rapid iteration.

  • Collaborate with researchers and engineers to build scalable infrastructure.

  • Publish and share learnings through internal documentation, open-source libraries, or technical reports that advance the field of scalable AI infrastructure.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

Preferred qualifications - we encourage you to apply if you meet some but not all of these:

  • Past experience working on distributed training for the world’s largest models to make them stable, reliable, and performant.

  • Track record of improving research productivity through infrastructure design or process improvements.

  • Contributions to open-source ML infrastructure such as PyTorch, XLA, Megatron-LM, or DeepSpeed.

Logistics

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$175k – $250k per year • Remote/Hybrid • Full-Time • 4+ years exp • Master's Degree • New York
C++
Python
SQL
C++
PyTorch C++
TensorFlow C++
AI/ML
PyTorch
Scikit-learn
TensorFlow
Apply
AI / ML Engineer 6 hours ago
$45k – $98k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad
AI/ML
PyTorch
TensorFlow
Apply
$215k – $393k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • High School Diploma • San Francisco • Mountain View
AI/ML
JAX
TensorFlow
TPU
Apply
$32k – $77k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Moscow
Python
Databases
FAISS
AI/ML
Hadoop
Hugging Face
LangChain
LLM
NLP
PyTorch
smolagents
Spark
Apply
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco
Python
Rust
TypeScript
JavaScript
AI/ML
Fine-tuning
Frontend
React.js
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
Fine-tuning
LoRA
PEFT
DevOps
CI/CD
Kubernetes
SRE
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
CUDA
CUDA Toolkit
NCCL
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
$180k – $260k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco
Python
AI/ML
ChatGPT
OpenAI
OpenAI Codex
Cybersecurity
FedRAMP
Apply
$180k – $210k per year • Equity • In office • Full-Time • San Francisco
Node JS
JavaScript
Databases
PostgreSQL
DevOps
PagerDuty
Web3
TRM Labs
Management
Slack
Apply
$252k – $335k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco
AI/ML
ChatGPT
Human-in-the-Loop
OpenAI
OpenAI Codex
DevOps
SLI/SLO/SLA
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.