368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$200k – $475k per year
Location
Remote/Hybrid (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Thinking Machines Lab is an artificial intelligence research and product company based in San Francisco and founded in 2025. The company develops multimodal AI systems and open-weights models, such as Inkling, alongside developer tools like Tinker for model fine-tuning. It operates as a public benefit corporation focused on human-AI collaboration and open science, supported by significant venture capital investment.

The mission of Thinking Machines is to build AI that extends human will and judgment.

About the Role

We’re looking for an infrastructure engineer to own and evolve the security infrastructure that underpins our foundation models. In this role, you’ll work across compute, storage, networking, and data platforms, making sure our systems are secure, reliable, and built to scale. You’ll shape controls, architecture, and tooling so that security is part of how the platform works by default. You’ll partner closely with research and product teams, enabling them to move quickly while keeping our models, data, and environments protected.

Note: This is an "evergreen role" that we keep open on an on-going basis to express interest. We receive many applications, and there may not always be an immediate role that aligns perfectly with your experience and skills. Still, we encourage you to apply. We continuously review applications and reach out to applicants as new opportunities open. You are welcome to reapply if you get more experience, but please avoid applying more than once every 6 months. You may also find that we put up postings for singular roles for separate, project or team specific needs. In those cases, you're welcome to apply directly in addition to an evergreen role.

What You’ll Do

  • Architect security patterns for platforms and services, including network segmentation, service-to-service authentication, RBAC, and policy enforcement in Kubernetes and cloud environments.

  • Manage identity, access, and secrets for humans and services: workload and cross-cloud identity, least-privilege IAM, and secrets management.

  • Build secure platforms for data ingestion, processing, and curation: classification, encryption, access controls, and safe sharing patterns across teams.

  • Write threat models and review designs with researchers and engineers to help them ship features and experiments in a safe, scalable way.

  • Automate security checks and build guardrails: policy-as-code, secure infrastructure baselines, validation in CI/CD, and tools that make the secure path the easiest one.

Skills and Qualifications

Minimum qualifications:

  • Bachelor’s degree or equivalent experience in engineering, or similar.

  • Strong background with containers and orchestration (e.g., Kubernetes) and how to secure them (namespaces, network policies, pod security, admission controls, etc.)

  • Practical experience with Infrastructure as Code (Terraform or similar), including secure patterns for provisioning networks, IAM, and shared services.

  • Solid understanding of cloud networking and security: VPCs, load balancers, service discovery, mTLS, firewalls, and zero-trust-style architectures.

  • Proficiency with a systems language such as Rust and scripting in Python for building platform components and internal tools.

  • Evidence of owning complex, production-critical systems, including debugging issues that span infra, security, and application layers.

Preferred qualifications - we encourage you to apply if you meet some even if you don't meet all of these:

  • Experience with ML infrastructure, GPU clusters, or large-scale training environments (schedulers, job queues, shared storage, multi-tenant clusters).

  • Background in AI labs, HPC environments, or ML-heavy organizations where both security and performance are first-class concerns.

  • Experience profiling and tuning high-throughput systems, and an ability to reason about the cost of additional security layers.

  • Talks, blogs, or publications on infrastructure security, distributed systems, or performance engineering.

  • Open-source contributions to security, orchestration, observability, or infrastructure tooling.

  • Familiarity with securing specialized hardware (GPUs, TPUs) and their integrations into training and inference pipelines.

Logistics

  • Location: This role is based in San Francisco, California.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $200,000 - $475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$61k – $143k per year (Estimated) • Remote/Hybrid • Full-Time • 4+ years exp • Cambridge
Go
Node JS
Python
TypeScript
JavaScript
Databases
ElasticSearch
MySQL
Frontend
GraphQL
Next.js
React.js
DevOps
CI/CD
Kubernetes
Apply
up to $48k per year (net) • Remote/Hybrid • Moscow
Node JS
Python
TypeScript
JavaScript
Node JS
Nest.JS
Python
Django
FastAPI
Databases
Apache Kafka
pgvector
Pinecone
PostgreSQL
Qdrant
RabbitMQ
Redis
AI/ML
Chain-of-Thought
Claude
Claude Code
Copilot
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
RAG
Anthropic
Function Calling
OpenAI
Structured Outputs
Frontend
GraphQL
Next.js
React.js
Redux
Redux Toolkit
Zustand
Mobile
State Management
DevOps
AWS
CI/CD
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Kubernetes
Prometheus
Yandex Cloud
GitHub
GitLab
Apply
$24k – $64k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Chennai
Python
Scala
SQL
TypeScript
JavaScript
Java
Python
pySpark
Java
Spring Boot
Databases
Apache Kafka
Databricks
AI/ML
AI Agents
Copilot
Google ADK
LLM
NLP
Prompt Engineering
Spark
Devin
Model Context Protocol
Frontend
Angular
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
OpenShift
GitHub
Analytics
ETL/ELT
Apply
$23k – $62k per year (Estimated) • In office • Full-Time • 10+ years exp • Gurgaon
C#
SQL
Databases
Apache Kafka
Redis
DevOps
Azure
CI/CD
Docker
gRPC
Kubernetes
Apply
Backend Engineer 1 day ago
$45k – $57k per year • In office • Full-Time • 3+ years exp • Tokyo
Python
TypeScript
JavaScript
Python
FastAPI
Databases
PostgreSQL
Frontend
Next.js
React.js
DevOps
Amazon EC2
AWS
AWS CDK
CI/CD
Docker
GitHub Actions
Vercel
Amazon CloudWatch
Amazon S3
GitHub
IAM
Management
Linear
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 5+ years exp • San Francisco
Apply
$300k – $475k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco
Python
Rust
TypeScript
JavaScript
AI/ML
Fine-tuning
Frontend
React.js
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
Fine-tuning
LoRA
PEFT
DevOps
CI/CD
Kubernetes
SRE
Apply
$350k – $475k per year • In office • Full-Time • 4+ years exp • San Francisco • New York
C++
Python
C++
PyTorch C++
AI/ML
PyTorch
Ray
Reinforcement Learning
RLHF
DPO
InfiniBand
NCCL
Post-training
PPO
TPU
DevOps
Kubernetes
SLURM
SRE
Apply
$350k – $475k per year • In office • Full-Time • San Francisco • New York
AI/ML
CUDA
CUDA Toolkit
NCCL
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 2 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.