Overview
News
Salaries
Products
People
Growth
Offices
Financials

Overview

Simple, illustrated, math-heavy guides to complex topics. The Illustrated Guide to LLM Inference at Scale (GPU memory, FSDP, Ulysses, Ring Attention) and Zero to Quantum (Grover.

News

Report
News darshanfofadiya.com 7 months ago
TurboQuant: Google
A deep technical breakdown of Google
Read more
Report
News darshanfofadiya.com 7 months ago
Zero to Quantum: From Curiosity to Writing Your Own Quantum Algorithms
A visual, code-first series that takes you from knowing nothing about quantum computing to understanding and designing quantum algorithms. Every concept explained with math, code, and no logical leaps.
Read more
Report
News darshanfofadiya.com 8 months ago
Ulysses Sequence Parallelism for 1M Context LLM Inference
How Ulysses sequence parallelism splits 1M tokens across GPUs for long-context LLM inference. All-to-All communication, comparison with tensor parallelism. Step-by-step calculations.
Read more
Report
News darshanfofadiya.com 8 months ago
GPU Memory for LLM Inference: Why Llama-70B Doesn
Why running Llama-70B with 1 million token context doesn
Read more
Report