866,360open jobs
54,486companies
146,147added this week
Browse all
Salary
≈ $166k – $358k per year (Estimated)
Location
Hybrid (San Francisco, New York, United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 28, 2026. First seen by Alion on Sep 1, 2026.

Overview
Company
Impact
Profile match
General Compute is the neocloud for SambaNova, Cerebras, Positron and d-Matrix. Prefill on GPUs, decode on purpose-built silicon — dedicated racks, one contract, one set of SLAs.

About Us

General Compute is the neocloud for alternative chips.

Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware - we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5-7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.

We closed a $15M seed round in May 2026, and have since closed a $400M debt facility - $100M funded upfront by Upper90, with the balance available for drawdown - collateralized by our inference chips.

About the role

You'll take a new model and get it running - correctly - on our ASIC in record time. When a frontier model drops, the only question that matters is how fast we can land it on our silicon and start serving it. You own that loop: from reference weights, through the compiler, to first correct tokens. The low-level runtime is co-owned with our hardware partner today; your job is everything it takes to get a brand-new architecture compiled, verified, and fast on top of it.

The bet of this role is that bring-up should be an agentic loop, not a hand-port. You'll build the harness of agents that compiles, runs, diffs against reference, and localizes failures - so the marginal model comes up faster than the last one did. Correctness first, optimization second: get it right, prove it's right, then make it cheap. This is a senior IC role on a small team. You'll own the bring-up pipeline, not tickets.

What You'll Do:

  • Own model bringup end-to-end. Take a new architecture - a frontier LLM, an MoE, a multimodal model - from reference weights to first correct tokens running on our ASIC, in days, not quarters.

  • Build the agentic bringup loop. The differentiator isn't hand-porting one model - it's the harness of agents that compiles, runs, diffs against reference, localizes the failing op, and iterates without you in the inner loop. Each model you land should make the loop better at landing the next one.

  • Live in the compiler. Graph capture, IR lowering, op coverage, kernel selection - when a model won't compile or produces wrong numbers, the fix is yours, whether it's a missing lowering, a fused-kernel bug, or a numerics mismatch.

  • Own correctness before speed. Build the verification harness - layer-by-layer activation diffs, logit parity, end-to-end evals - that proves a freshly brought-up model matches reference before anyone trusts a token of it.

  • Then optimize. Once it's correct, make it fast: operator fusion, quantization, memory layout, batching and KV-cache behavior on our hardware. Bringup gets it running; this is where it earns its cost-per-token.

  • Work shoulder-to-shoulder with our hardware partner's compiler and runtime team. You're the person who turns 'the chip can technically run this' into 'this model is live and correct in production.

What we need from you:

  • 5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.

  • You've taken a model architecture you didn't design and made it run - and run correctly - on a target it wasn't written for. Numerics debugging doesn't scare you.

  • Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching. You can read a new model's reference implementation and know what will be hard to lower.

  • Comfortable inside a compiler stack - MLIR/LLVM, XLA, or a vendor graph compiler - at the level of IR, lowering, and op coverage, not just calling into one.

  • Fluent with agentic tooling. You'd rather build the agent that runs the tedious bringup loop than run it by hand - and you have the taste to know where the loop still needs a human.

  • Self-directed. We don't assign tickets - you'll see the next model coming and have it half brought-up before anyone asks.

Nice-to-haves:

  • Have worked on a non-NVIDIA accelerator - TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar - at the compiler or model-bringup layer.

  • Kernel-level experience in CUDA, Triton, or a vendor kernel language. You know why a fused attention kernel beats three unfused ops.

  • Have built eval and numerics-verification harnesses (logit parity, activation diffing) for models in production.

  • Contributed to a graph compiler or serving runtime - XLA, TVM, MLIR, vLLM, TGI, TensorRT-LLM, or SGLang.

  • Have built agent loops or LLM-driven tooling that did real engineering work, not demos.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
866,360 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
$286k – $327k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • San Francisco • McLean • Cambridge • San Jose • New York
Python
Go
Java
C#
C++
Scala
C++
PyTorch C++
AI/ML
CUDA Toolkit
AI Agents
PyTorch
LLM
CUDA
Hugging Face
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$160k – $200k per year • In office • Full-Time • Bachelor's Degree • Emeryville
Python
AI/ML
JAX
Multimodal AI
Diffusion Models
Accelerate
TensorFlow
PyTorch
Machine Learning
Apply
$132k – $158k per year • In office • 3+ years exp • PhD • New York
AI/ML
JAX
Multimodal AI
PyTorch
Machine Learning
Apply
$153k – $198k per year • Remote (United States) • Full-Time • 5+ years exp
Python
SQL
Databases
Redis
AI/ML
Spark
dbt
Scikit-learn
TensorFlow
PyTorch
Amazon SageMaker
Feature Store
Recommender Systems
Machine Learning
DevOps
AWS
Apply
≈ $136k – $288k per year (Estimated) • Hybrid • 5+ years exp • Bachelor's Degree • Atlanta
Python
SQL
Databases
PostgreSQL
Databricks
Milvus
pgvector
Pinecone
OpenSearch
AI/ML
LangGraph
LangChain
LlamaIndex
Model Context Protocol
Fine-tuning
Scikit-learn
AI Agents
NLP
spaCy
AWS Bedrock
Transformers
Pandas
NumPy
PyTorch
LLM
RAG
NLTK
Machine Learning
DevOps
Terraform
Datadog
CI/CD
Git
AWS
Docker
Kubernetes
AWS Lambda
Amazon S3
Amazon ECS
Analytics
Matplotlib
Apply
≈ $40k – $95k per year (Estimated) • In office • Full-Time • 10+ years exp • Master's Degree • Gurgaon
AI/ML
Fine-tuning
AI Agents
Transformers
TensorFlow
PyTorch
LLM
RAG
BERT
Hugging Face
DevOps
Docker
Kubernetes
Apply
$11k – $17k per year • Remote (likely EAEU) • 1+ year exp • Moscow
Python
SQL
Python
SQLAlchemy
FastAPI
Asyncio
Celery
Pydantic
Alembic
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
LangGraph
LangChain
vLLM
LLM
OpenAI
LiveKit
Structured Outputs
DevOps
Rest API
Docker Compose
WebSockets
Git
Docker
Ubuntu
GitHub
Linux
QA
Pytest
Apply
Scrum Master 4C 9 hours ago
Hybrid • Full-Time • Bachelor's Degree • Gurgaon
AI/ML
AI Agents
Edge AI
DevOps
Azure
Management
Agile
Scrum
Apply
≈ $21k – $44k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Bengaluru
Java
AI/ML
AI Agents
Edge AI
Apply
Tech Lead - DevOps 4C 9 hours ago
≈ $20k – $48k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Bengaluru
AI/ML
AI Agents
Edge AI
DevOps
Azure DevOps
Chef
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Bamboo
GitHub
GitLab
Linux
Apply
≈ $180k – $326k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • San Francisco • New York
AI/ML
Groq
vLLM
SGLang
TensorRT
TensorRT-LLM
TGI
LLM
Cerebras
TPU
AWS Trainium
Speculative Decoding
KV Cache
Apply
≈ $137k – $301k per year (Estimated) • Hybrid • Full-Time • San Francisco • New York
AI/ML
Quantization
Cerebras
TPU
KV Cache
DevOps
ZooKeeper
Nomad
SLURM
Kubernetes
Apply
≈ $197k – $366k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $178k – $330k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
Cerebras
CoreWeave
Apply
≈ $212k – $395k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • San Francisco
AI/ML
vLLM
SGLang
TensorRT
TensorRT-LLM
TGI
OpenRouter
Cerebras
TPU
InfiniBand
DevOps
Kubernetes
Platform Engineering
HPC
Apply
$90k – $98k per year • In office • Internship • Bachelor's Degree • San Francisco
Apply
$83k – $88k per year • In office • Bachelor's Degree • San Francisco
Apply
$90k – $98k per year • In office • Internship • Bachelor's Degree • San Francisco
Apply
$78k – $86k per year • In office • Internship • Bachelor's Degree • San Francisco
Apply
$85k – $97k per year • In office • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
866,360 more open roles from verified company boards, updated every day.