368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$240k – $280k per year
Location
In office (San Francisco)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

About the Role

We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle - turning racks of GPUs into running inference clusters without a human touching a runbook. The Research and Inference team is your customer: today they file tickets and wait; the target state is that they issue a single API call to stand up, scale, or tear down a cluster, and the system takes care of the rest. The platform is manifest-driven such that teams declare the desired state of a cluster or host - shape, topology, software stack - and the system is responsible for reconciling reality to that manifest, continuously, through every stage of its lifecycle. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry a piece of hardware or a cluster from one state to the next-taking it from bare metal to a fully functioning AI cluster for training or inference.

You'll write production code which is typed, tested, versioned, and deployed through CI/CD that models infrastructure state and reconciles it, the same way a Kubernetes controller reconciles a cluster's desired state. Success looks like eliminating manual provisioning work, not documenting it better.

A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production.

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call - no ticket, no human in the loop.
  • Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
  • Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
  • Partner with the inference/ML platform team: understand the cluster shapes they need - topology, interconnect, scheduling constraints - and encode them as first-class abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code - this is a product, not a collection of Ansible playbooks.

Requirements

Core requirements (all levels):

  • Strong software engineering background in Go, Python, Rust, or similar - you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems - designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Nice to have:

  • Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
  • Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
  • Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $240,000 - $280,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$220k – $325k per year • Remote/Hybrid • Full-Time • 15+ years exp • New York
DevOps
AWS
Azure
CI/CD
GCP
Kubernetes
Platform Engineering
Cybersecurity
Threat Modeling
Apply
$140k – $253k per year (Estimated) • In office • 5+ years exp • Long Beach
C++
Rust
C++
Protobuf
AI/ML
Human-in-the-Loop
DevOps
CI/CD
Vector
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$16k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$36k – $140k per year (Estimated) • In office • Amsterdam
Go
Python
Rust
Databases
Apache Kafka
NATS
AI/ML
CUDA
CUDA Toolkit
Together AI
Human-in-the-Loop
InfiniBand
NCCL
DevOps
Ansible
CI/CD
Kubernetes
Self-Healing
Apply
$30k – $78k per year (Estimated) • In office • 5+ years exp
Go
Python
Rust
TypeScript
Databases
Apache Kafka
NATS
AI/ML
AI Agents
LLM
RAG
Semantic Search
Together AI
Function Calling
Knowledge Graph
Semantic Search
DevOps
ArgoCD
GitOps
Grafana
Incident Management
Kubernetes
Prometheus
Management
Slack
Apply
$140k – $170k per year • Remote • Full-Time • 3+ years exp • San Francisco
SQL
AI/ML
Claude
Claude Code
Together AI
Apply
$200k – $250k per year • Remote • Full-Time • 5+ years exp • San Francisco
Python
SQL
AI/ML
Together AI
InfiniBand
DevOps
HPC
Apply
$93k – $219k per year (Estimated) • In office • London
Go
Python
Rust
Databases
Apache Kafka
NATS
AI/ML
CUDA
CUDA Toolkit
Together AI
Human-in-the-Loop
InfiniBand
NCCL
DevOps
Ansible
CI/CD
Kubernetes
Self-Healing
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Databases
Databricks
Google BigQuery
SAP HANA
Snowflake
AI/ML
Knowledge Graph
DevOps
Azure
Apply
$70k – $196k per year • Remote/Hybrid • Full-Time • 5+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
DevOps
SLI/SLO/SLA
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$293k – $385k per year • In office • Full-Time • San Francisco
AI/ML
OpenAI
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.