368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$190k – $270k per year
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

Build the infrastructure powering the next generation of AI.

At Together AI, you’ll build and operate one of the world’s largest GPU fleets used for frontier model training and inference. This isn’t a traditional infrastructure role-we’re looking for engineers who love building systems, automating everything, and solving problems at massive scale.

You’ll Thrive Here If You:

  • Love building systems that replace repetitive operational work.
  • Think of infrastructure as a software engineering problem.
  • Enjoy solving hard problems with no existing playbook.
  • Care deeply about performance, reliability, and scale.
  • Want to build technology that powers frontier AI models.

Our mission is simple: build AI infrastructure that largely runs itself-where intelligent systems deploy, monitor, diagnose, optimize, and heal GPU fleets at massive scale. Every system you build will directly improve the speed, efficiency, and reliability of one of the world’s most advanced AI compute platforms.

If you enjoy writing software more than clicking dashboards, obsess over eliminating manual work, and want to build infrastructure that manages tens of thousands of GPUs autonomously, we’d love to talk.

Responsibilities

  • Design and build fleet automation systems that provision, validate, deploy, upgrade, repair, and retire GPU clusters with minimal human intervention.
  • Build AI Infrastructure Agents that automate deployment, root-cause failures, incident triage, and autonomous remediation.
  • Develop Fleet Intelligence platforms that continuously monitor hardware health, firmware, networking, storage, thermals, and workload performance to predict failures before they impact customers.
  • Build software that maximizes GPU availability, utilization, performance, and reliability across thousands of accelerators.
  • Create automated validation systems for GPUs, InfiniBand/RoCE fabrics, NVLink/NVSwitch, storage, and distributed AI workloads.
  • Build internal platforms and developer tools that allow infrastructure to be managed through software-not manual operations.
  • Continuously improve deployment velocity, reliability, and operational efficiency through automation.
  • Partner closely with hardware, networking, platform, and AI teams to push the limits of AI infrastructure.

Requirements

  • 3+ years building distributed systems, infrastructure platforms, or large-scale backend software.
  • Strong software engineering skills in Python, Go, or Rust.
  • Experience building platforms, automation systems, or developer infrastructure.
  • Experience with Linux, Kubernetes, Terraform, Ansible, or similar infrastructure technologies.
  • Strong systems thinking with the ability to understand problems across hardware and software.
  • A passion for solving complex infrastructure challenges through software.
  • An automation-first mindset -if a task is repeated, your instinct is to build a system to eliminate it.

Bonus Experience

  • GPU infrastructure, CUDA, NCCL, NVLink/NVSwitch
  • InfiniBand or RoCE networking
  • Bare-metal provisioning and lifecycle management
  • Large-scale AI training or inference clusters
  • Hardware health monitoring and predictive failure detection
  • Distributed storage systems
  • AI agents and autonomous infrastructure operations

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance and other competitive benefits. The US base salary range for this full-time position is: $190,000 - $270,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Sr DevOps Engineer 1 day ago
$23k – $59k per year (Estimated) • In office • Full-Time
AI/ML
AI Agents
DevOps
Ansible
CI/CD
Configuration Management
SLI/SLO/SLA
Terraform
Apply
$55k – $157k per year (Estimated) • Remote • Full-Time • Sydney
C++
Go
Lua
Python
C++
CMake
Databases
ActiveMQ
Aerospike
Apache Kafka
Cassandra
DevOps
Docker
gRPC
Kubernetes
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$77k – $215k per year (Estimated) • In office • Contractor • 3+ years exp • Singapore
JavaScript
Python
TypeScript
Python
Django
FastAPI
Databases
PostgreSQL
Redis
AI/ML
AI Agents
LangGraph
LLM
Prompt Engineering
LangChain
Edge AI
Frontend
Angular
React.js
Vue.js
Apply
$18k – $46k per year (Estimated) • In office • Full-Time • Moscow
DevOps
Ansible
ArgoCD
AWS
CI/CD
Docker
GitLab CI
Helm
Kubernetes
OpenTofu
Terraform
Terragrunt
Yandex Cloud
GitLab
Apply
$120k – $150k per year • In office • Full-Time • 1+ year exp • Bachelor's Degree • San Francisco
Python
SQL
Databases
Snowflake
AI/ML
dbt
Together AI
Marketing
Amplitude
Salesforce
Apply
$36k – $140k per year (Estimated) • In office • Amsterdam
Go
Python
Rust
Databases
Apache Kafka
NATS
AI/ML
CUDA
CUDA Toolkit
Together AI
Human-in-the-Loop
InfiniBand
NCCL
DevOps
Ansible
CI/CD
Kubernetes
Self-Healing
Apply
$30k – $78k per year (Estimated) • In office • 5+ years exp
Go
Python
Rust
TypeScript
Databases
Apache Kafka
NATS
AI/ML
AI Agents
LLM
RAG
Semantic Search
Together AI
Function Calling
Knowledge Graph
Semantic Search
DevOps
ArgoCD
GitOps
Grafana
Incident Management
Kubernetes
Prometheus
Management
Slack
Apply
$140k – $170k per year • Remote • Full-Time • 3+ years exp • San Francisco
SQL
AI/ML
Claude
Claude Code
Together AI
Apply
$200k – $250k per year • Remote • Full-Time • 5+ years exp • San Francisco
Python
SQL
AI/ML
Together AI
InfiniBand
DevOps
HPC
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 3 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.