791,964open jobs
50,470companies
123,351added this week
Browse all
Salary
$260k – $300k per year
Location
Remote (United States)
Seniority
Staff · 7+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 23, 2026. Together AI scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

About the Role

Together AI is building the AI Native Cloud, an end-to-end platform for the full generative AI lifecycle, combining the fastest LLM inference engine with state-of-the-art GPU cloud infrastructure. The Together Cloud team builds the [Together GPU Clusters](https://www.together.ai/gpu-clusters) flagship IaaS product that provides high-performance, AI-ready GPU clusters through a self-serve cloud console, along with the virtualized infrastructure layer powering Together's inference, RL, and fine-tuning products.

As a Staff Software Engineer focusing on AI Compute in the Together Cloud org, you will set technical direction for and build major components of the next generation AI cloud platform - a highly available, global cloud infrastructure with cutting-edge virtualization of the latest ML hardware: GB300s/VRs, BlueField DPUs, InfiniBand and dual/quad-plane RoCEv2 fabrics. That virtualized computing platform powers our own SaaS products - inference, RL, and fine-tuning - and serves external cloud customers through self-serve offerings such as on-demand/reserved Kubernetes/Slurm clusters, across dozens of data centers and hundreds of thousands of GPUs.

This is an architect-and-build role. Fully automated bootstrapping of GPU data centers, high-performance virtualization of GPU compute and DC networking without compromising isolation or portability, and fault-tolerant decentralized control planes - you'll set the architecture for these across our global and in-DC services, and be a key owner of the hardest parts, in the code as well as the design. Your designs will span the IaaS layer of a greenfield Vera Rubin data center up to the global management plane that schedules capacity across all of them. At this level the job is as much leverage as code: the standards you set and the engineers you grow decide how fast the rest of Together Cloud ships.

Responsibilities

  • Own the GPU and network virtualization stack: the hypervisor, kernel, and SDN work that keeps GPU compute and DC networking high-performance, portable, and strongly isolated across heterogeneous hardware.
  • Own the in-DC IaaS layer: architect and roadmap the services, Kubernetes operators, and libraries that provision and manage compute, storage, and networks in our data centers - VMs, parallel filesystems, VPCs, and InfiniBand partitions; lead its build-out for a new Vera Rubin data center with thousands of GPUs, from hardware bring-up to customer-facing API.
  • Design the GPU scheduling and global management plane: the distributed control plane behind on-demand and reserved clusters across dozens of data centers, including the systems that scale per-cluster limits and automate the onboarding of new capacity.
  • Architect monitoring and automated remediation for fault tolerance: the strategy for automated detection, isolation, and recovery of failed nodes that keeps distributed pretraining and large-scale inference running through hardware failures.
  • Set technical direction across teams: lead design reviews, resolve cross-cutting architectural disagreements, unblock cross-team dependencies and integration risks, and define the standards other engineers build against - measured in cluster reliability, time-to-first-GPU on new capacity, and quality at scale.
  • Grow the team: mentor senior and junior engineers, deepen the team's expertise in virtualization, DC networking, and GPU infrastructure, and help raise the hiring bar for Together Cloud.
  • Set the engineering bar: create the testing frameworks, tools, and developer documentation that make our systems robust and usable by other teams, and shape the core, open-source Together AI platform.

To be successful you'll need to be deeply technical and an excellent communicator - expert software development fundamentals, deep systems knowledge and troubleshooting instincts, and the leadership and diplomacy skills to align teams that don't report to you. Much of this work starts ambiguous, and we expect you to define the scope yourself and drive it to production.

Requirements

  • 7+ years of professional software development experience, with expert-level proficiency in at least one backend language (Golang desired), writing high-performance, well-tested, production-quality code.
  • Track record of owning the architecture of large distributed systems from blank page to production at scale, including the judgment calls that could not be reversed cheaply.
  • Deep experience building and operating globally distributed, high-performance microservice architectures across one or more cloud providers (AWS, Azure, GCP).
  • Expert systems knowledge across compute, networking, and storage - including concurrency, memory management, performant I/O, and scale at a global level.
  • Demonstrated technical leadership beyond your own commits: mentoring senior engineers, leading design reviews, and driving alignment across teams that do not report to you.
  • Excellent communication and diplomacy skills - able to write design docs that settle arguments, and to work effectively with technical and non-technical stakeholders.
  • Experience building and operating reliable, customer-facing production systems at scale, and owning the infrastructure automation (Terraform, Ansible), observability (Prometheus, Grafana), and CI/CD (GitHub Actions, ArgoCD) that keep them healthy.

Preferred Qualifications (not must haves)

  • Deep Kubernetes internals experience, such as implementing non-trivial Kubernetes operators, device/storage/network plugins, custom schedulers, or patches to Kubernetes itself
  • Deep experience with VMs/hypervisors, such as QEMU/KVM, cloud-hypervisor, VFIO, virtio, PCIE passthrough, Kubevirt, SR-IOV
  • Deep experience with DC networking tech + solutions, such as VLAN, VXLAN, VPN, VPC, OVS/OVN
  • Experience with Cluster API or similar
  • Experience working on high-performance compute, networking, and/or storage
  • Experience virtualizing GPUs and/or InfiniBand
  • Experience building IaaS or PaaS systems at scale
  • Experience with DPUs/SmartNICs
  • GPU programming, NCCL, CUDA knowledge

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers in our journey in building the next generation AI infrastructure.

Compensation

We offer competitive compensation, startup equity, health insurance, and other benefits, as well as flexibility in terms of remote work. The US base salary range for this full-time position is: $260,000 - $300,000 + equity + benefits. Our salary ranges are determined by location, level and role. Individual compensation will be determined by experience, skills, and job-related knowledge.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at  https://www.together.ai/privacy  

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
791,964 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
$69k – $101k per year • Equity • Remote (India) • Full-Time • 5+ years exp • Bachelor's Degree
Java
Java
Maven
Gradle
Databases
ClickHouse
Apache Kafka
DevOps
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Management
Agile
Apply
Software Engineer 5 months ago
≈ $141k – $232k per year (Estimated) • Remote (United States, EST hours) • Full-Time • 5+ years exp
Python
JavaScript
Python
Celery
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
LangChain
Embeddings
OCR
Frontend
Next.js
React.js
DevOps
Terraform
Vercel
AWS
Apply
≈ $31k – $77k per year (Estimated) • Remote (Romania) • Bucharest
JavaScript
Rust
Node JS
Databases
PostgreSQL
Redis
DynamoDB
Frontend
Redux
React.js
DevOps
Terraform
GitHub Actions
AWS CDK
AWS
Docker
Kubernetes
AWS Lambda
Web3
Anchor
Solana
DeFi
Helius
Metaplex
Apply
$243k – $295k per year • Equity • Hybrid • 5+ years exp • Bachelor's Degree • San Mateo
Python
Go
C#
AI/ML
Multimodal AI
Kubeflow
DevOps
Amazon S3
Apply
≈ $125k – $231k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Austin
Go
JavaScript
Java
TypeScript
AI/ML
LLM
Frontend
React.js
Mobile
State Management
DevOps
Rest API
CI/CD
Kubernetes
Apply
$69k – $131k per year • In office • Secret • Full-Time • 2+ years exp • Bachelor's Degree • Concord
Python
PowerShell
Bash
DevOps
Terraform
Ansible
VMWare
CI/CD
Docker
Kubernetes
Grafana
Platform Engineering
KVM
Linux
Management
Confluence
Agile
Scrum
Kanban
Apply
$120k – $250k per year • Equity 0.5–2.5% • Remote (United States) • Full-Time • 3+ years exp • San Francisco
Python
Go
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
Weaviate
Milvus
DevOps
GCP
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
Management
Stripe
Apply
≈ $18k – $43k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Pune
JavaScript
DevOps
Rest API
GCP
Azure
AWS
Platform Engineering
SOAP
Management
ServiceNow
ITIL
Microsoft Office
Apply
$57k – $109k per year • In office • TS/SCI • Full-Time • Bachelor's Degree • West Valley City
Python
C++
DevOps
Puppet
Ansible
Chef
Git
Docker
Kubernetes
Configuration Management
Linux
Management
Confluence
Jira
Agile
Apply
$80k – $200k per year • Equity 0.1–2.5% • In office • Full-Time • New York
AI/ML
World Models
DevOps
GCP
Azure
AWS
Robotics
Imitation Learning
Apply
$150k – $160k per year • In office • Full-Time • Bachelor's Degree • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
Machine Learning
DevOps
Git
Platform Engineering
Apply
In office • Internship • Bachelor's Degree • San Francisco
Python
TypeScript
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
DevOps
Platform Engineering
Apply
In office • Internship • Bachelor's Degree • San Francisco
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
DevOps
Git
Platform Engineering
Apply
≈ $68k – $146k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Bengaluru
Python
Databases
MinIO
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Together AI
Pre-training
InfiniBand
DevOps
Terraform
Ansible
Helm
Prometheus
GitOps
ArgoCD
Kubernetes
Grafana
Thanos
Amazon S3
HPC
Linux
Apply
$250k – $300k per year • Remote (United States) • Full-Time • 5+ years exp • San Francisco
Python
Go
Rust
TypeScript
Databases
NATS
Apache Kafka
AI/ML
Cursor
ElevenLabs
Fine-tuning
Reinforcement Learning
Function Calling
AI Agents
LLM
RAG
Semantic Search
Together AI
Pre-training
Semantic Search
Knowledge Graph
Tool Use
DevOps
Prometheus
GitOps
ArgoCD
Kubernetes
Grafana
Incident Management
Management
Slack
Apply
Founding Engineer 12 hours ago
$110k – $180k per year • Equity 0.1–1% • In office • Full-Time • San Francisco
Python
JavaScript
Node JS
AI/ML
Vertex AI
OpenAI
Anthropic
Frontend
Next.js
React.js
DevOps
Azure
Kubernetes
Apply
$60k – $84k per year • In office • Internship • San Francisco
Python
JavaScript
TypeScript
AI/ML
Copilot
Cursor
Claude
Claude Code
Model Context Protocol
Vertex AI
AI Agents
LLM
OpenAI
Anthropic
LLM Guardrails
Tool Use
Frontend
Next.js
React.js
DevOps
Azure
AWS
Kubernetes
Apply
Founding Engineer 12 hours ago
$120k – $200k per year • Equity 0.5–1% • In office • Full-Time • 3+ years exp • San Francisco
TypeScript
AI/ML
Claude
AI Agents
LLM
DevOps
GitHub
Management
Slack
Apply
Founding Engineer 12 hours ago
$120k – $180k per year • Equity 0.1–0.6% • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Rust
C++
Assembly
Assembly
Binary Ninja
AI/ML
Machine Learning
DevOps
QEMU
Linux
Cybersecurity
Ghidra
IDA Pro
Apply
$80k – $120k per year • In office • Internship • Bachelor's Degree • San Francisco
Python
C++
AI/ML
Multimodal AI
Machine Learning
Apply
See all jobs
This is one of many
791,964 more open roles from verified company boards, updated every day.