961,875open jobs
58,127companies
159,282added this week
Browse all
Salary
≈ $175k – $332k per year (Estimated)
Location
Hybrid (San Francisco, New York, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 29, 2026. First seen by Alion on Sep 14, 2026.

Overview
Company
Impact
Profile match
General Compute is the neocloud for SambaNova, Cerebras, Positron and d-Matrix. Prefill on GPUs, decode on purpose-built silicon — dedicated racks, one contract, one set of SLAs.

About us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware - we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5-7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility - $100M funded upfront by Upper90, with the balance available for drawdown - collateralized by our inference chips.

About the Role

Getting a model correct and fast on our silicon is only half the problem - the other half is serving it. You'll build and own the inference layer that sits between a bought-up model and a live customer request: request scheduling, batching, KV-cache management, autoscaling across our ASIC fleet, and the failure modes that only show up at real traffic and real scale.

This is a founding role on a small team, which means the scope is wide and the ownership is real: there's no separate SRE org to hand reliability to and no platform team to hand infra to. You'll design the serving architecture, then be the person paged when it breaks. The bet is that a serving stack built specifically for our hardware - not adapted from a GPU-first framework - is a durable edge, and you're the person who proves that out in production.

What You'll Do:

  • Own the inference serving stack end-to-end. Design and build the system that takes a bring-up-verified model and serves it in production: request routing, batching, scheduling, and autoscaling a single model's serving replicas.

  • Push cost-per-token down. Continuously tune batching strategy, KV-cache handling, and hardware utilization to widen the throughput advantage over GPU-based serving.

  • Build for reliability from day one. Put in place the monitoring, alerting, and failover that make a fast-moving inference stack trustworthy under real customer load - and be the one who responds when it isn't.

  • Work at the boundary with the compiler and bring-up team. Define the interface between "a model is correct and compiled" and "a model is live and fast," and push issues back to the right side of that line.

  • Shape the roadmap, not just the backlog. As a founding engineer, you'll help decide what we build next in serving - multi-tenant isolation, speculative decoding, new scheduling strategies - not just execute a spec someone else wrote.

  • Set the technical bar for the team you're helping build. Early architecture and code-quality decisions you make here will shape how the serving team operates as it grows.

What We Need From You:

  • 5+ years building and operating production systems at the infrastructure layer, ideally including a high-throughput or low-latency serving system.

  • Direct experience with LLM inference serving - request batching, KV-cache management, continuous batching, or similar - in a production environment, not just research code.

  • Comfortable owning reliability: you've been on call for a system that mattered, and you design for failure rather than reacting to it after the fact.

  • Strong systems fundamentals - concurrency, networking, scheduling - deep enough to reason about performance at the hardware level, not just the application level.

  • Self-directed and comfortable with ambiguity. This is a founding role: there's no existing playbook to follow, and you'll help write it.

Nice-to-Haves:

  • Experience serving models on non-NVIDIA accelerators (TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar).

  • Familiarity with serving frameworks such as vLLM, TGI, TensorRT-LLM, or SGLang, and an opinion on where they fall short.

  • Experience running infrastructure at a small company or in a founding/early-engineer capacity before.

  • Exposure to capacity planning or fleet management for specialized hardware.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
961,875 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$130k – $147k per year • Hybrid • Full-Time • New York
Python
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Spark
Model Context Protocol
MLFlow
AI Agents
LLM
Agentic Workflows
Machine Learning
DevOps
CI/CD
Platform Engineering
Incident Management
Apply
≈ $165k – $340k per year (Estimated) • In office • 8+ years exp • Sunnyvale
Python
SQL
C#
ABAP
ABAP
SAP BTP
AI/ML
Copilot
LangGraph
LangChain
LlamaIndex
Model Context Protocol
AI Agents
RAG
A2A
Multi-Agent Systems
Copilot Studio
DevOps
CI/CD
Apply
$80k – $160k per year • In office • 8+ years exp • Bachelor's Degree • Atlanta
AI/ML
Copilot
AI Agents
Edge AI
DevOps
GCP
Azure
AWS
Apply
Azure AI Engineer 7 months ago
$103k – $135k per year • In office • Full-Time • 5+ years exp • Master's Degree • Huntington Beach
Python
Databases
Databricks
Delta Lake
Microsoft Fabric
AI/ML
Model Context Protocol
Fine-tuning
Scikit-learn
Prompt Engineering
AI Agents
Prophet
TensorFlow
PyTorch
LLM
Time Series Forecasting
OpenAI
A2A
Multi-Agent Systems
Copilot Studio
Machine Learning
DevOps
Azure
CI/CD
Apply
≈ $130k – $246k per year (Estimated) • In office • Full-Time • 2+ years exp • Atlanta
TypeScript
SQL
Elixir
Databases
PostgreSQL
SQLite
AI/ML
Claude
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
LLM
Anthropic
Human-in-the-Loop
Tool Use
Vercel AI SDK
DevOps
GCP
OpenTelemetry
Datadog
WebSockets
AWS
Cloudflare
Apply
Software Engineer 3 days ago
≈ $107k – $198k per year (Estimated) • Equity • Remote (United States) • 3+ years exp
Python
TypeScript
SQL
AI/ML
Cursor
Claude
Claude Code
AI Agents
LLM
Apply
Founding Engineer 3 days ago
$150k – $220k per year • Remote (United States) • Full-Time • 4+ years exp • New York
Python
AI/ML
Fine-tuning
Reinforcement Learning
AI Agents
LLM
Robotics
Reinforcement Learning
Apply
≈ $37k – $63k per year (Estimated) • In office • Internship • Chicago
AI/ML
LLM
Apply
≈ $96k – $241k per year (Estimated) • In office • 6+ years exp • Master's Degree • Kigali
AI/ML
Fine-tuning
Reinforcement Learning
AI Agents
LLM
Explainable AI
Machine Learning
Apply
$180k – $400k per year • Equity • In office • Full-Time • 2+ years exp • San Francisco
Python
JavaScript
TypeScript
AI/ML
LangGraph
LangChain
Embeddings
AI Agents
CrewAI
LLM
RAG
Multi-Agent Systems
Tool Use
Frontend
React.js
DevOps
GCP
CI/CD
AWS
Apply
≈ $166k – $361k per year (Estimated) • Hybrid • Full-Time • San Francisco • New York
AI/ML
Groq
vLLM
CUDA Toolkit
Quantization
Multimodal AI
AI Agents
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
Mixture of Experts
Cerebras
CUDA
Triton
TPU
AWS Trainium
MLIR
Apache TVM
XLA
KV Cache
Apply
≈ $134k – $299k per year (Estimated) • Hybrid • Full-Time • San Francisco • New York
AI/ML
Quantization
Cerebras
TPU
KV Cache
DevOps
ZooKeeper
Nomad
SLURM
Kubernetes
Apply
≈ $192k – $362k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $172k – $326k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $208k – $391k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
vLLM
SGLang
TensorRT
TensorRT-LLM
TGI
OpenRouter
Cerebras
TPU
InfiniBand
DevOps
Kubernetes
Platform Engineering
HPC
Apply
$208k – $340k per year • Hybrid • Full-Time • 8+ years exp • San Francisco • New York
Databases
Databricks
AI/ML
Embeddings
AI Agents
LLM
Apply
$119k – $239k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Atlanta • Chicago • New York • San Francisco • Los Angeles
Apply
$55k – $58k per year • In office • San Francisco
Apply
≈ $59k – $122k per year (Estimated) • Remote (United States) • San Francisco
Management
Microsoft Office
Apply
$127k – $269k per year • Hybrid • Full-Time • 5+ years exp • San Francisco
Design
Figma
Apply
See all jobs
This is one of many
961,875 more open roles from verified company boards, updated every day.