Overview
News
Technologies
Salaries
Products
People
Growth
Financials

Overview

Doubleword runs high-volume async and batch AI inference up to 90% cheaper than real-time APIs. OpenAI-compatible, for agents, evals and pipelines.

News

News Doubleword 21 day ago
On-the-fly snapshot compression for elastic inference at scale
Accelerating snapshot-based model serving with on-the-fly LZ4 memory compression in CRIU.
Read more
Report
News Doubleword 21 day ago
The case for disaggregated LLM serving
Three checkable conditions under which disaggregated prefill is never worse than aggregated serving, and the reasons it's usually better. Plus an inventory-theory model of traffic drift: cold starts, safety stock, and a map of the shocks a deployment can
Read more
Report
News Doubleword 1 month ago
You Could Have Come Up With Kimi Delta Attention
Guide to the DeltaNet Family of linear attention mechanisms.
Read more
Report
News Doubleword 1 month ago
NVLink, NVSwitch, and all that
Why scale-up links are fast and short, how NVSwitch grew the domain from a board to a rack, and what TPU, UALink, and scale-up Ethernet do differently.
Read more
Report
News Doubleword 1 month ago
Reverse-engineering NVIDIA's cuda-checkpoint for faster cold starts
Freezing a live CUDA process to host memory and thawing it again, what the driver does - and doesn't - do to make that work, and how understanding that lets us restore CUDA processes up to 4x faster.
Read more
Report
Pro access
Upgrade to see all 15 mentions
Upgrade to a paid plan to read every media mention of this company - funding news, awards, product launches and press releases from all the outlets writing about it.
Every media mention and press release
Funding news, awards and product launches
Fresh coverage from every outlet writing about the company
Upgrade now
Cancel anytime. Secure checkout. Instant activation.