824,647open jobs
53,159companies
134,960added this week
Browse all
Salary
$150k – $300k per year
Location
In office (San Francisco)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 27, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
GPU orchestration, model inference, and an agent platform. Adopt Paladin, Ion, or Talos independently, or connect the full Cumulus stack.

About the role

Cumulus Labs builds the software that turns raw GPU capacity into fast, cheap, production AI. We're looking for an ML Platforms Engineer to help build and run the orchestration layer underneath our inference and agent products, the system that schedules workloads, allocates GPUs, and keeps a heterogeneous, multi-cloud fleet running at high utilization.

We care more about how you think than which languages are on your resume. Our stack today includes Go, Kubernetes, and Terraform, but we're looking for someone who can walk into any part of a production system, understand it, and make it better, not someone who only knows one toolchain.

What you'll do

  • Build and extend our GPU orchestrator: scheduling, fractional allocation, live workload migration across GPUs with no downtime

  • Design and evolve multi-tenant primitives: quotas, isolation, usage metering, a tenant-facing inference gateway

  • Own observability for the fleet: metrics, logs, and traces at scale

  • Debug hard, systems-level problems across the stack, from scheduling logic down to GPU memory and networking

  • Make real architectural decisions, not just implement someone else's design

  • Ship fast, own your systems end to end, and work directly with the founder

What we're looking for

  • Excellent fundamentals: data structures, algorithms, distributed systems concepts, and the judgment to apply the right pattern to the right problem

  • Real production experience, ideally with systems that had to stay up and scale under load

  • Strong design instincts: you can reason about tradeoffs, not just follow a framework's conventions

  • Fast learner who can go deep in unfamiliar territory; specific experience with Go or Kubernetes is a plus, not a requirement

  • Comfortable using modern AI coding tools (we use Claude Code heavily) to move fast without losing rigor

  • You want to work in person, in a small team, solving problems nobody has solved before

Why Cumulus

We're small, early, and building the systems layer for the next generation of AI infrastructure. You'll own real infrastructure from day one, not tickets in a backlog.

1. Intro call with a founder (30 min) - background, mutual fit

2. Technical deep dive (30-45 min) - a real problem from our orchestrator, discussed live

3. Take-home or paired session on a scoped systems problem

4. References

We move fast: most candidates hear back within a few days at each stage, and we aim to get from first contact to offer in under two weeks.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,647 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$94k – $294k per year • Hybrid • Full-Time • 12+ years exp • Associate's Degree • Dallas • Tampa • Atlanta • Columbus • Houston
Python
Java
Databases
Databricks
Delta Lake
AI/ML
LangGraph
AutoGen
LangChain
Claude
Spark
MLFlow
Vertex AI
AI Agents
CrewAI
LLM
RAG
OpenAI
Anthropic
LLMOps
LLM Evaluation
Agentic Workflows
Multi-Agent Systems
DevOps
CI/CD
Apply
$74k – $262k per year • Hybrid • Full-Time • 12+ years exp • Associate's Degree • Dallas • Tampa • Atlanta • Columbus • Houston
Python
Java
Databases
Databricks
Delta Lake
AI/ML
LangChain
Claude
Spark
DSPy
MLFlow
Vertex AI
AI Agents
LLM
RAG
OpenAI
Hugging Face
LLMOps
Context Engineering
LLM Evaluation
Agentic Workflows
Multi-Agent Systems
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Apply
≈ $140k – $245k per year (Estimated) • Remote (United States) • Public Trust • Full-Time • Bachelor's Degree • Salt Lake City
Python
SQL
Python
pySpark
Databases
Databricks
Delta Lake
Microsoft Fabric
AI/ML
Spark
MLFlow
XGBoost
Scikit-learn
PyTorch
RAG
Machine Learning
DevOps
Terraform
Azure
CI/CD
Git
Cybersecurity
HIPAA
FedRAMP
Microsoft Entra ID
Analytics
Power BI
ETL/ELT
SSIS
Azure Data Factory
SSAS
Apply
≈ $145k – $263k per year (Estimated) • In office • Austin
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
$120k – $170k per year • In office • Chicago
JavaScript
C++
AI/ML
CUDA Toolkit
LLM
CUDA
Frontend
WebGPU
DevOps
HPC
Game Dev
GLSL
Apply
Backend Engineer 1 day ago
≈ $56k – $147k per year (Estimated) • Remote (Turkey) • Full-Time • 5+ years exp • Bachelor's Degree • Istanbul
Java
Kotlin
SQL
Java
Spring Boot
Databases
RabbitMQ
Apache Kafka
DevOps
Rest API
CI/CD
Git
Docker
Kubernetes
Management
Agile
Apply
$60k – $180k per year • Equity 0.5–2% • Remote (United States) • Full-Time • 1+ year exp • San Francisco
Python
TypeScript
SQL
Python
FastAPI
AI/ML
AI Agents
Edge AI
Browser Agents
Frontend
Next.js
Mobile
Clean Architecture
DevOps
GitHub Actions
Docker
Kubernetes
Linux
Apply
$150k – $250k per year • Equity 0.3–0.5% • Remote (United States) • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco
Python
Databases
FAISS
ElasticSearch
AI/ML
Fine-tuning
AI Agents
NLP
LLM
RAG
OpenAI
Hugging Face
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Marketing
X (Twitter)
Apply
$133k – $302k per year • In office • Full-Time • 12+ years exp • Chicago • Milwaukee • Herndon • Dallas • Columbus
AI/ML
Prompt Engineering
AI Agents
LLM
RAG
Human-in-the-Loop
Knowledge Graph
Agentic Workflows
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Apply
Dev-Ops Engineer 1 day ago
$31k – $37k per year • In office • Full-Time • 3+ years exp • Bengaluru
DevOps
Terraform
CI/CD
Git
AWS
Docker
Kubernetes
AWS Lambda
Amazon EC2
Amazon S3
Linux
Marketing
Shopify
Apply
$130k – $175k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Apply
≈ $105k – $207k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$140k – $295k per year • In office • 5+ years exp • San Francisco
Apply
$100k – $125k per year • In office • 2+ years exp • San Francisco
Python
Apply
$145k – $235k per year • In office • 3+ years exp • San Francisco
Apply
See all jobs
This is one of many
824,647 more open roles from verified company boards, updated every day.