Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
Blog. Three checkable conditions under which disaggregated prefill is never worsethan aggregated serving, and the reasons it's usually better. Plus aninventory-theory model of traffic drift: cold starts, safety stock, and amap of the shocks a deployment can ride out.
News
The case for disaggregated LLM serving
Three checkable conditions under which disaggregated prefill is never worse than aggregated serving, and the reasons it's usually better. Plus an inventory-theory model of traffic drift: cold starts, safety stock, and a map of the shocks a deployment can
Read more
Report
Throughputmaxxing: DeepSeek-V4-Flash on Isambard-AI
Single-node serving throughput for DeepSeek-V4-Flash: baseline, MLA/DP-attention, and MoE kernel choice.
Read more
Report
NVLink, NVSwitch, and all that
Why scale-up links are fast and short, how NVSwitch grew the domain from a board to a rack, and what TPU, UALink, and scale-up Ethernet do differently.
Read more
Report
Reverse-engineering NVIDIA's cuda-checkpoint for faster cold starts
Freezing a live CUDA process to host memory and thawing it again, what the driver does - and doesn't - do to make that work, and how understanding that lets us restore CUDA processes up to 4x faster.
Read more
Report
Width vs. depth: speculating on the margin
Some thinking about how to trade off batching and speculative decoding in a running inference engine.
Read more
Report
Pro access
Upgrade to see all 34 mentions
Upgrade to a paid plan to read every media mention of this company - funding news, awards, product launches and press releases from all the outlets writing about it.
Every media mention and press release
Funding news, awards and product launches
Fresh coverage from every outlet writing about the company
Upgrade now
Cancel anytime. Secure checkout. Instant activation.
