Confirmed on the employer's own hiring board on Sep 28, 2026. First seen by Alion on Sep 1, 2026.
About Us
General Compute is the neocloud for alternative chips.
Inference is fragmenting: purpose-built silicon from SambaNova, Cerebras, Positron, d-Matrix, and others already beats GPUs on decode, and we productionize that hardware - we buy the racks, find the data center space, and run it for our customers. Each piece of hardware runs the workload it's actually built for: prefill stays on GPUs, decode moves to the chip built for it, and today that means generating tokens 5-7× faster than existing GPU-based competitors. Our customers are frontier labs, fast-growing AI application companies, and asset-light clouds.
We closed a $15M seed round in May 2026, and have since closed a $400M debt facility - $100M funded upfront by Upper90, with the balance available for drawdown - collateralized by our inference chips.
About the role
You'll take a new model and get it running - correctly - on our ASIC in record time. When a frontier model drops, the only question that matters is how fast we can land it on our silicon and start serving it. You own that loop: from reference weights, through the compiler, to first correct tokens. The low-level runtime is co-owned with our hardware partner today; your job is everything it takes to get a brand-new architecture compiled, verified, and fast on top of it.
The bet of this role is that bring-up should be an agentic loop, not a hand-port. You'll build the harness of agents that compiles, runs, diffs against reference, and localizes failures - so the marginal model comes up faster than the last one did. Correctness first, optimization second: get it right, prove it's right, then make it cheap. This is a senior IC role on a small team. You'll own the bring-up pipeline, not tickets.
What You'll Do:
Own model bringup end-to-end. Take a new architecture - a frontier LLM, an MoE, a multimodal model - from reference weights to first correct tokens running on our ASIC, in days, not quarters.
Build the agentic bringup loop. The differentiator isn't hand-porting one model - it's the harness of agents that compiles, runs, diffs against reference, localizes the failing op, and iterates without you in the inner loop. Each model you land should make the loop better at landing the next one.
Live in the compiler. Graph capture, IR lowering, op coverage, kernel selection - when a model won't compile or produces wrong numbers, the fix is yours, whether it's a missing lowering, a fused-kernel bug, or a numerics mismatch.
Own correctness before speed. Build the verification harness - layer-by-layer activation diffs, logit parity, end-to-end evals - that proves a freshly brought-up model matches reference before anyone trusts a token of it.
Then optimize. Once it's correct, make it fast: operator fusion, quantization, memory layout, batching and KV-cache behavior on our hardware. Bringup gets it running; this is where it earns its cost-per-token.
Work shoulder-to-shoulder with our hardware partner's compiler and runtime team. You're the person who turns 'the chip can technically run this' into 'this model is live and correct in production.
What we need from you:
5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.
You've taken a model architecture you didn't design and made it run - and run correctly - on a target it wasn't written for. Numerics debugging doesn't scare you.
Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching. You can read a new model's reference implementation and know what will be hard to lower.
Comfortable inside a compiler stack - MLIR/LLVM, XLA, or a vendor graph compiler - at the level of IR, lowering, and op coverage, not just calling into one.
Fluent with agentic tooling. You'd rather build the agent that runs the tedious bringup loop than run it by hand - and you have the taste to know where the loop still needs a human.
Self-directed. We don't assign tickets - you'll see the next model coming and have it half brought-up before anyone asks.
Nice-to-haves:
Have worked on a non-NVIDIA accelerator - TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar - at the compiler or model-bringup layer.
Kernel-level experience in CUDA, Triton, or a vendor kernel language. You know why a fused attention kernel beats three unfused ops.
Have built eval and numerics-verification harnesses (logit parity, activation diffing) for models in production.
Contributed to a graph compiler or serving runtime - XLA, TVM, MLIR, vLLM, TGI, TensorRT-LLM, or SGLang.
Have built agent loops or LLM-driven tooling that did real engineering work, not demos.

