Overview
News
Technologies
Salaries
Products
People
Growth
Financials

Overview

RunInfra helps teams build AI applications and inference infrastructure through chat, with model selection, GPU benchmarking, optimization, and deployment in one workflow. Backed by Y Combinator.

News

News RunInfra 15 days ago
Model Library Model APIs
Browse RunInfra's hosted model library for published pricing, limits, capabilities, and availability. DeepSeek V4 Flash: API access is available. Nemotron 3.5 Lightning 30B: API access is available. Qwen3.8 2.4T A95B: API access is available. Qwen3.8 27B:
Read more
Report
News RunInfra 26 days ago
With software alone, one B200 beats the LPU and gets close to Cerebras
One rented B200 served gpt-oss-120b at 411 tok/s on stock SGLang and 1,366 tok/s after four settings, with single prompts at 2,215. That is 3.5x the fastest GPU provider on Artificial Analysis, past Groq and SambaNova, at 72 percent of Cerebras. No kernel
Read more
Report
Pro access
Upgrade to see all 7 mentions
Upgrade to a paid plan to read every media mention of this company - funding news, awards, product launches and press releases from all the outlets writing about it.
Every media mention and press release
Funding news, awards and product launches
Fresh coverage from every outlet writing about the company
Upgrade now
Cancel anytime. Secure checkout. Instant activation.