368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$350k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure research engineer to design, optimize, and maintain the compute foundations that power large-scale language model training. You will develop high-performance ML kernels (e.g., CUDA, CuTe, Triton), enable efficient low-precision arithmetic, and improve the distributed compute stack that makes training large models possible.

This role is perfect for an engineer who enjoys working close to the metal and across the research boundary. You’ll collaborate with researchers and systems architects to bridge algorithmic design with hardware efficiency. You’ll prototype new kernel implementations, profile performance across hardware generations, and help define the numerical and parallelism strategies that determine how we scale next-generation AI systems.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Design and implement custom ML kernels (e.g., CUDA, CuTe, Triton) for core LLM operations such as attention, matrix multiplication, gating, and normalization, optimized for modern GPU and accelerator architectures.

  • Design and think through compute primitives to reduce memory bandwidth bottlenecks and improve kernel compute efficiency.

  • Collaborate with research teams to align kernel-level optimizations with model architecture and algorithmic goals.

  • Develop and maintain a library of reusable kernels and performance benchmarks that serve as the foundation for internal model training.

  • Contribute to infrastructure stability and scalability, ensuring reproducibility, consistency across precision formats, and high utilization of compute resources.

  • Document and share insights through internal talks, technical papers, or open-source contributions to strengthen the broader ML systems community.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.

  • Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases

  • Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.

  • Thrive in a highly collaborative environment involving many, different cross-functional partners and subject matter experts.

  • A bias for action with a mindset to take initiative to work across different stacks and different teams where you spot the opportunity to make sure something ships.

  • Proficiency in CUDA, CuTe, Triton, or other GPU programming frameworks.

  • Demonstrated ability to analyze, profile, and optimize compute-intensive workloads.

Preferred qualifications - we encourage you to apply if you meet some but not all of these:

  • Experience training or supporting large-scale language models with tens of billions of parameters or more.

  • Track record of improving research productivity through infrastructure design or process improvements.

  • Experience developing or tuning kernels for deep learning frameworks such as PyTorch, JAX, or custom accelerators.

  • Familiarity with tensor parallelism, pipeline parallelism, or distributed data processing frameworks.

  • Experience implementing low-precision formats (FP8, INT8, block floating point) or contributing to related compiler stacks (e.g., XLA, TVM).

  • Contributions to open-source GPU, ML systems, or compiler optimization projects.

  • Prior research or engineering experience in numerical optimization, communication-efficient training, or scalable AI infrastructure.

Logistics

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
$35k – $86k per year (Estimated) • In office • Moscow
Python
SQL
Databases
Apache Kafka
AI/ML
LLM
Model Context Protocol
RAG
DevOps
Grafana
Apply
$18k – $51k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
ClickHouse
PostgreSQL
AI/ML
Feature Store
Hadoop
LLM
DevOps
Ansible
CI/CD
Docker
GitLab
Kubernetes
Apply
$71k – $154k per year (Estimated) • In office • Full-Time • Dublin
Java
Databases
Apache Kafka
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
AWS
CI/CD
Kubernetes
GitHub
Apply
Lead AI Engineer 1 day ago
$30k – $73k per year (Estimated) • In office • Full-Time • Pune
Python
AI/ML
Fine-tuning
LLM
Reinforcement Learning
LLM Guardrails
AI Agents
DevOps
CI/CD
Docker
GitOps
Helm
Kubernetes
OpenShift
Platform Engineering
Vector
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco
Python
Rust
TypeScript
JavaScript
AI/ML
Fine-tuning
Frontend
React.js
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
Fine-tuning
LoRA
PEFT
DevOps
CI/CD
Kubernetes
SRE
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
CUDA
CUDA Toolkit
NCCL
Apply
$170k – $220k per year • Equity 1–2.8% • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
Python
Django
AI/ML
AI Agents
Context Engineering
LLM
LLM Evaluation
RAG
Apply
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
Senior ML Engineer 1 hour ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • San Francisco
Go
JavaScript
Ruby
Scala
Apply
$360k – $530k per year • In office • Full-Time • Bachelor's Degree • San Francisco
MATLAB
Python
MATLAB
Simulink
AI/ML
OpenAI
Robotics
Digital Twin
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.