368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$36k – $140k per year (Estimated)
Location
In office (Amsterdam)
Seniority
Junior
Overview
Company
Impact
Profile match
Together AI (Together Computer, Inc.) is a full-stack AI infrastructure and cloud platform headquartered in San Francisco, California. Founded in 2022 by prominent AI researchers and system engineers - including CEO Vipul Ved Prakash, CTO Ce Zhang, Chief Scientist Tri Dao (co-creator of FlashAttention), Chris Ré, and Percy Liang - the company operates as an "AI Native Cloud" designed to train, fine-tune, and deploy open-source generative AI models at scale with high performance and optimized unit economics.

About the Role

We're looking for a Software Engineer to build the systems that treat infrastructure as software. This role owns the software state machines that provision hardware, bring it into service, and manage its full lifecycle - turning racks of GPUs into running inference clusters without a human touching a runbook. The Research and Inference team is your customer: today they file tickets and wait; the target state is that they issue a single API call to stand up, scale, or tear down a cluster, and the system takes care of the rest. The platform is manifest-driven such that teams declare the desired state of a cluster or host - shape, topology, software stack - and the system is responsible for reconciling reality to that manifest, continuously, through every stage of its lifecycle. You will design the engines that manifest the schema, the engines that execute against it, and the workflows that carry a piece of hardware or a cluster from one state to the next-taking it from bare metal to a fully functioning AI cluster for training or inference.

You'll write production code which is typed, tested, versioned, and deployed through CI/CD that models infrastructure state and reconciles it, the same way a Kubernetes controller reconciles a cluster's desired state. Success looks like eliminating manual provisioning work, not documenting it better.

A product mindset - you've built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.You build it, you own it. You are not only responsible for delivering the software but also for operating and supporting it in production.

Responsibilities

  • Build the provisioning state machine: design and implement the software that models the full lifecycle of a physical host from discovery, inference bring-up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned states and transitions.
  • Build the self-service API: design declarative APIs and a control plane so the inference team can request, scale, and tear down inference clusters with one API call - no ticket, no human in the loop.
  • Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or replacement, and reintroduce healthy capacity into the pool automatically.
  • Own reliability of the pipeline: idempotency, retries, rollback, and drift detection so the provisioning system is as dependable as any other production service.
  • Partner with the inference/ML platform team: understand the cluster shapes they need - topology, interconnect, scheduling constraints - and encode them as first-class abstractions in the platform.
  • Engineer it like software: strong typing, automated tests, code review, versioning, and CI/CD for infrastructure code - this is a product, not a collection of Ansible playbooks.

Requirements

Core requirements (all levels):

  • Strong software engineering background in Go, Python, Rust, or similar - you write and test real software for a living.
  • Experience with durable workflow orchestration tools such as Temporal, Cadence, or equivalent to run long-lived, manifest-driven workflows that survive failures and resume mid-execution.
  • Experience building software control planes or orchestration systems that model state and reconcile it over time (e.g., Kubernetes controllers/operators, custom reconciliation loops, workflow engines).
  • Experience with event-driven systems - designing and building software around message queues, event streams, or pub/sub (e.g., Kafka, NATS, SQS) rather than polling or cron-driven scripts.
  • A product mindset. You’ve built internal platforms or APIs consumed by other engineering teams and care about the developer experience of what you ship.

Nice to have:

  • Exposure to bare-metal provisioning (PXE/iPXE, Redfish/IPMI, BMC) and/or networking fundamentals (VLANs, BGP, fabric design), or GPU/accelerator infrastructure.
  • Experience with GPU cluster software stacks (NCCL, CUDA, InfiniBand/RoCE).
  • Prior work at a hyperscaler, GPU cloud, or datacenter-scale infrastructure organization.
  • Systems programming in Rust or Go.

About Together AI

Together AI is a research-driven artificial intelligence company. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models. We have contributed to leading open-source research, models, and datasets to advance the frontier of AI, and our team has been behind technological advancement such as FlashAttention, Hyena, FlexGen, and RedPajama. We invite you to join a passionate group of researchers and engineers in our journey in building the next generation AI infrastructure.

Equal Opportunity

Together AI is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.

Please see our privacy policy at  https://www.together.ai/privacy.  

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Amsterdam
$121k – $147k per year • Remote • Full-Time • 5+ years exp • PhD • Columbus
Java
Java
Spring Boot
Databases
Apache Kafka
DevOps
Azure
Azure AKS
Azure DevOps
Bicep
CI/CD
Dynatrace
GitLab
Kubernetes
Splunk
Terraform
Management
Jira
Apply
AI Engineer 1 hour ago
$120k – $130k per year • In office • Full-Time • Texas
Python
SQL
Databases
Apache Kafka
Snowflake
AI/ML
Amazon SageMaker
Hadoop
DevOps
CI/CD
Git
GitLab
Analytics
ETL/ELT
Apply
$29k – $73k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Guadalajara
Python
DevOps
Amazon CloudWatch
Azure
CI/CD
Datadog
Docker
Kubernetes
Prometheus
Terraform
Apply
$38k – $91k per year (Estimated) • Remote • Full-Time • 8+ years exp • PhD • Guadalajara
PowerShell
Python
Databases
Amazon Aurora
DynamoDB
AI/ML
Amazon SageMaker
AWS Bedrock
AWS Bedrock AgentCore
Ray
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
Datadog
FinOps
GCP
Git
GitLab
GitLab CI
IAM
Jenkins
JFrog Artifactory
Kubernetes
New Relic
Service Mesh
Splunk
Terraform
Cybersecurity
HIPAA
ISO 27001
PCI DSS
SOC 2
Apply
$222k – $277k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco
DevOps
CI/CD
Immutable Infrastructure
Apply
$30k – $78k per year (Estimated) • In office • 5+ years exp
Go
Python
Rust
TypeScript
Databases
Apache Kafka
NATS
AI/ML
AI Agents
LLM
RAG
Semantic Search
Together AI
Function Calling
Knowledge Graph
Semantic Search
DevOps
ArgoCD
GitOps
Grafana
Incident Management
Kubernetes
Prometheus
Management
Slack
Apply
$140k – $170k per year • Remote • Full-Time • 3+ years exp • San Francisco
SQL
AI/ML
Claude
Claude Code
Together AI
Apply
$200k – $250k per year • Remote • Full-Time • 5+ years exp • San Francisco
Python
SQL
AI/ML
Together AI
InfiniBand
DevOps
HPC
Apply
$93k – $219k per year (Estimated) • In office • London
Go
Python
Rust
Databases
Apache Kafka
NATS
AI/ML
CUDA
CUDA Toolkit
Together AI
Human-in-the-Loop
InfiniBand
NCCL
DevOps
Ansible
CI/CD
Kubernetes
Self-Healing
Apply
$40k – $95k per year (Estimated) • In office
Go
Python
Rust
Databases
Apache Kafka
NATS
AI/ML
CUDA
CUDA Toolkit
Together AI
Human-in-the-Loop
InfiniBand
NCCL
DevOps
Ansible
CI/CD
Kubernetes
Self-Healing
Apply
$92k – $223k per year (Estimated) • In office • Contractor • Amsterdam
Databases
PostgreSQL
Redis
AI/ML
AI Agents
Gemini
Google ADK
LangChain
LlamaIndex
Prompt Engineering
Vertex AI
DevOps
GCP
Kubernetes
Management
Slack
Apply
$105k – $252k per year (Estimated) • In office • 8+ years exp • Amsterdam
SQL
Apply
$91k – $216k per year (Estimated) • In office • Full-Time • 10+ years exp • Amsterdam
Java
JavaScript
Java
Gradle
Frontend
npm
pnpm
DevOps
ArgoCD
AWS
Bazel
CI/CD
Datadog
GitHub Actions
GitOps
Helm
JFrog Artifactory
Kubernetes
SRE
Terraform
GitHub
IAM
Cybersecurity
GDPR
Apply
$101k – $242k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Amsterdam
Python
Databases
Databricks
AI/ML
Spark
Apply
$67k – $169k per year (Estimated) • In office • 4+ years exp • Amsterdam
Python
SQL
Python
pySpark
AI/ML
Airflow
Spark
Marketing
Salesforce
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.