824,647open jobs
53,159companies
134,960added this week
Browse all
Salary
≈ $137k – $301k per year (Estimated)
Location
Hybrid (San Francisco, New York, United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 14, 2026.

Overview
Company
Impact
Profile match
General Compute is the neocloud for SambaNova, Cerebras, Positron and d-Matrix. Prefill on GPUs, decode on purpose-built silicon — dedicated racks, one contract, one set of SLAs.

About us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware - we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5-7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility - $100M funded upfront by Upper90, with the balance available for drawdown - collateralized by our inference chips.

About the role

You will build the inference cloud itself - the control plane, API, and serving layer that turn racks into a sellable product. There's no existing platform team to inherit or manage, no legacy system to work around, and no established playbook to follow - just the platform itself to build, with reliability treated as core infrastructure from day one rather than something bolted on after the first outage.The technical problem is also genuinely unsolved elsewhere. The fleet is heterogeneous by design - GPUs for prefill, multiple ASIC vendors for decode - so there's no single-vendor playbook to lean on; you'll be defining how a mixed-hardware inference cloud gets scheduled, routed, and served reliably, in close partnership with the teams standing up the physical fleet.

What you'll do:

  • Build/own the control plane - routing, model placement, scheduling across a mixed ASIC/GPU pool

  • Build the API and serving layer exposing rack capacity as a sellable product

  • Build in reliability and observability from day one

  • Scale the platform ahead of the demand curve

  • Partner closely with data center deployment and model bring-up teams

  • Be a founding technical voice on platform architecture

  • Work at the boundary with the inference team. Own fleet-wide routing, placement, and capacity; hand off to the Inference Engineer for per-model serving performance (batching, KV-cache, quantization) rather than owning that layer yourself

What we need from you:

  • Strong systems engineering background on distributed cloud control planes and multi-tenant orchestration at scale (e.g., Kubernetes-style schedulers)

  • Comfort being a high-impact IC rather than a manager

  • Track record building reliability from scratch

  • Experience with resource scheduling or bin-packing algorithms for heterogeneous compute pools

  • Experience with multi-cluster federation or distributed consensus systems (Raft, ZooKeeper, or similar)

  • Comfort with hardware heterogeneity/ambiguity

  • Genuine interest in being an early hire at a ~6-7 person company

Nice-to-haves:

  • Experience with distributed job schedulers or cluster orchestration systems (Kubernetes, Nomad, Slurm, or similar)

  • Experience running non-NVIDIA accelerators (TPUs/ASICs) in production

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
824,647 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
San Francisco
$185k – $258k per year • In office • Full-Time • 8+ years exp • Washington • London
DevOps
Azure
Platform Engineering
Management
ServiceNow
Agile
Apply
≈ $84k – $175k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Georgetown
Python
SQL
C#
DevOps
Linux
Windows
Apply
$200k – $300k per year • In office • Full-Time • New York
Management
Stripe
Apply
≈ $115k – $212k per year (Estimated) • Remote (United States) • Full-Time • 6+ years exp
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
AI/ML
Prompt Engineering
AI Agents
Frontend
React.js
Apply
$85k – $168k per year • Remote (United States) • 2+ years exp • Bachelor's Degree • Washington
Python
PowerShell
DevOps
Puppet
Chef
Azure
CI/CD
Windows Server
Configuration Management
Hyper-V
Linux
Windows
Unix
DNS
Cybersecurity
Active Directory
Management
OneDrive
Apply
$66k – $106k per year • In office • 5+ years exp • London
Python
Python
FastAPI
Django
AI/ML
LangGraph
LangChain
LlamaIndex
Time Series Forecasting
Frontend
Redux
DevOps
Terraform
CI/CD
AWS
Docker
Kubernetes
Cybersecurity
GDPR
Apply
$90k – $100k per year • In office • 4+ years exp
Python
Java
PHP
TypeScript
SQL
Ruby
Databases
MySQL
PostgreSQL
DynamoDB
Apache Kafka
Amazon Aurora
AI/ML
Spark
DevOps
Terraform
Ansible
Chef
CloudFormation
Git
AWS
Docker
Kubernetes
Platform Engineering
Configuration Management
GitHub
Amazon Kinesis
Linux
Apply
AI Engineer 1 day ago
In office
Python
C++
Swift
C++
TensorFlow C++
PyTorch C++
AI/ML
OpenCV
Triton Inference Server
Quantization
Scikit-learn
Multimodal AI
Computer Vision
TensorRT
TensorFlow
PyTorch
LLM
Point Cloud Library
Synthetic Data
Triton
Amazon SageMaker
ONNX Runtime
Machine Learning
Mobile
Core ML
ARKit
Metal
DevOps
CI/CD
AWS
Wi-Fi
Robotics
Open3D
Sensor Fusion
Apply
Remote (Singapore) • PhD • Singapore
Python
AI/ML
AutoGen
Weights & Biases
LangChain
DeepSeek
MLFlow
Reinforcement Learning
Quantization
Multimodal AI
Knowledge Distillation
Computer Vision
Chain-of-Thought
AI Agents
Transformers
TensorFlow
PyTorch
RAG
TensorBoard
Ray
Mixture of Experts
Hybrid Search
OpenAI
Hugging Face
Constitutional AI
Multi-Agent Systems
Model Distillation
Reward Modeling
Machine Learning
Cybersecurity
Empire
Apply
≈ $149k – $246k per year (Estimated) • Remote (United States) • 5+ years exp • Austin
JavaScript
TypeScript
Node JS
Frontend
React.js
DevOps
AWS
Kubernetes
Cybersecurity
SOC 2
Apply
≈ $180k – $326k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • San Francisco • New York
AI/ML
Groq
vLLM
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
Cerebras
TPU
AWS Trainium
Speculative Decoding
KV Cache
Apply
≈ $166k – $358k per year (Estimated) • Hybrid • Full-Time • San Francisco • New York
AI/ML
Groq
vLLM
CUDA Toolkit
Quantization
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
Mixture of Experts
Cerebras
CUDA
Triton
TPU
AWS Trainium
MLIR
Apache TVM
XLA
KV Cache
Apply
≈ $197k – $366k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $178k – $330k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $212k – $395k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
vLLM
SGLang
TensorRT
TensorRT-LLM
TGI
OpenRouter
Cerebras
TPU
InfiniBand
DevOps
Kubernetes
Platform Engineering
HPC
Apply
$174k – $235k per year • Equity • In office • Full-Time • 6+ years exp • San Francisco
Python
Go
Java
Rust
Ruby
C
C#
C++
C
MPI
DevOps
AWS CDK
CloudFormation
CI/CD
AWS
Docker
Kubernetes
Ubuntu
AWS Lambda
Amazon EC2
Amazon S3
HPC
Linux
Unix
TCP/IP
DNS
DHCP
Apply
$130k – $175k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • San Francisco
Apply
≈ $105k – $207k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
Apply
$140k – $295k per year • In office • 5+ years exp • San Francisco
Apply
$100k – $125k per year • In office • 2+ years exp • San Francisco
Python
Apply
See all jobs
This is one of many
824,647 more open roles from verified company boards, updated every day.