1,280,349open jobs
74,275companies
212,107added this week
Browse all
Salary
$300k – $400k per year
Location
In office (San Francisco, New York)
Visa
Sponsorship offered in the posting · H-1B filings in 12 months: 30 · for this role: 7
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 6, 2026. First seen by Alion on Oct 2, 2026.

Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer, Inference to own the reliability, scale, and efficiency of the systems that serve our models to real users. Our research and inference teams push the limits of model performance and serving efficiency; this role makes sure those gains reach production safely and stay up - powering Tinker's live, multi-tenant serving and the products built on top of our models.

This is a production-facing systems role at the center of the company. You'll be the bridge between cutting-edge inference techniques and the day-to-day reality of serving real traffic: rollouts, capacity, incidents, and everything that keeps a fast-growing platform online.

What You'll Do

  • Operate and scale the production inference systems that serve live traffic, including Tinker's multi-tenant serving platform

  • Own the rollout process for new models, model versions, and inference optimizations, ensuring safe, incremental deployment to production

  • Build and improve observability, alerting, and capacity planning so the team can detect, diagnose, and resolve production issues quickly

  • Partner with inference and research teams to productionize new serving techniques without compromising reliability

  • Lead incident response for production inference issues, driving root cause analysis and durable fixes

  • Design for graceful degradation, failover, and redundancy so that serving stays resilient as usage grows

  • Manage capacity and cost tradeoffs for serving infrastructure as traffic and model sizes scale

Skills & Qualifications

Minimum Qualifications

  • Experience operating large-scale, latency-sensitive production systems

  • Proficiency in Python and Go or another systems language

  • Experience with observability, monitoring, and incident response for production services

  • Strong understanding of distributed systems and how they fail at scale

Preferred Qualifications

  • Experience running production inference for large language models or other large-scale ML systems

  • Experience with deployment and rollout systems, such as canarying, blue/green deploys, or feature flags

  • Experience with capacity planning and cost optimization for GPU or TPU infrastructure

  • Familiarity with inference-specific techniques, such as batching, caching, or quantization, and their operational implications

  • Comfortable being on-call and leading incident response for critical production systems

  • Comfortable working with high autonomy in a fast-changing, early-stage environment

Logistics

  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $300,000 - $400,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,280,349 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
$125k – $200k per year • Equity • Remote (United States) • 7+ years exp • Bachelor's Degree • Denver
JavaScript
TypeScript
Node JS
Databases
MySQL
PostgreSQL
Frontend
React.js
DevOps
AWS
Amazon Kinesis
Management
Agile
Scrum
Apply
≈ $139k – $234k per year (Estimated) • Remote (United States) • Full-Time • New York
Python
Java
Rust
TypeScript
AI/ML
AI Agents
DevOps
CI/CD
AWS
Apply
$176k – $308k per year • Equity • In office • Full-Time • 5+ years exp • Bachelor's Degree • Mountain View
AI/ML
Machine Learning
Management
Slack
ServiceNow
Apply
$172k – $301k per year • Equity • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • San Diego
Java
Kotlin
SQL
C++
Scala
Kotlin
Glide
Databases
Snowflake
Databricks
Apache Iceberg
Delta Lake
Apache Kafka
Google BigQuery
Amazon Redshift
Trino
BigQuery
AI/ML
Spark
AI Agents
DevOps
Kubernetes
Cybersecurity
Polaris
Analytics
ETL/ELT
Management
ServiceNow
Scrum
Apply
$191k – $334k per year • Equity • Hybrid • Full-Time • 12+ years exp • Santa Clara
JavaScript
TypeScript
AI/ML
AI Agents
Hallucination
Multi-Agent Systems
Management
ServiceNow
Apply
≈ $75k – $164k per year (Estimated) • In office • Full-Time • Vancouver
Python
C
C++
C
U-Boot
DevOps
RTOS
CI/CD
Git
Platform Engineering
Linux
TCP/IP
DHCP
VLAN
Chips/EDA
OpenOCD
IoT
MQTT
Apply
$86k – $138k per year • In office • Secret • 8+ years exp • Bachelor's Degree • Las Cruces
Python
C++
Fortran
Databases
PostgreSQL
MS SQL
AI/ML
RAG
DevOps
GitLab CI
CI/CD
Jenkins
GitHub
GitLab
Linux
Unix
Management
Confluence
Agile
Apply
≈ $46k – $104k per year (Estimated) • Remote (United States, Brazil) • Full-Time
Python
SQL
AI/ML
Copilot
Claude
ChatGPT
dbt
Analytics
A/B Testing
Apply
≈ $59k – $116k per year (Estimated) • In office • 3+ years exp • Dallas
Python
PowerShell
DevOps
Rest API
IAM
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Microsoft Entra ID
Active Directory
Management
SharePoint
Apply
≈ $22k – $41k per year (Estimated) • In office • 3+ years exp • Moscow
Python
SQL
Mobile
AppsFlyer SDK
Apply
$300k – $400k per year • In office • Full-Time • 5+ years exp • San Francisco • New York
Python
C++
AI/ML
TPU
Apply
$300k – $400k per year • In office • Full-Time • San Francisco • New York
Python
AI/ML
TPU
DevOps
Kubernetes
Apply
$300k – $400k per year • In office • Full-Time • San Francisco • New York
Python
Rust
Databases
Delta Lake
Apache Kafka
AI/ML
Spark
Airflow
dbt
Multimodal AI
LLM
Ray
DevOps
Terraform
Apply
$300k – $350k per year • In office • Full-Time • San Francisco • New York
Python
Rust
AI/ML
AI Agents
DevOps
Kubernetes
Linux
Cybersecurity
Threat Modeling
Apply
$350k – $475k per year • In office • Full-Time • San Francisco
Python
AI/ML
JAX
TensorFlow
PyTorch
Post-training
Machine Learning
Apply
Senior SWE 10 hours ago
$300k – $500k per year • In office • Full-Time • 3+ years exp • San Francisco
Python
AI/ML
Post-training
Apply
$250k – $339k per year • Remote (United States) • Full-Time • 10+ years exp • Master's Degree • Atlanta • Cambridge • San Francisco • Thousand Oaks
Apply
Senior Data Engineer 4 hours ago
$200k – $250k per year • In office • Full-Time • 6+ years exp • San Francisco
Python
SQL
Databases
Snowflake
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Groq
E2B
DevOps
AWS
AWS Lambda
Management
Zapier
Apply
In office • Internship • San Francisco • Munich
Python
JavaScript
TypeScript
Python
FastAPI
Pydantic
Databases
PostgreSQL
pgvector
AI/ML
Langfuse
Pydantic AI
LLM
LLM Guardrails
Frontend
React.js
Vite
TanStack Router
Chakra UI
Apply
Operating Engineer 5 hours ago
$159k per year • In office • Full-Time • 3+ years exp • San Francisco
Management
Outlook
Microsoft Office
Apply
See all jobs
This is one of many
1,280,349 more open roles from verified company boards, updated every day.