Overview
News
Technologies
Salaries
Products
People
Growth
Financials
Overview
I am Moncef Abboud. I work as a platform engineer and write here about subjects that interest me, including programming, systems, and LLMs & AI. I enjoy open source and have made a few contributions to various projects, which you can check out on my GitHub.
News
How Profitable is LLM Inference? Doing the Math on Kimi K3
A look at LLM inference economics (batch size, GPU count, and the Pareto frontier that sets token prices) applied to Kimi K3 with back-of-the-envelope math.
Read more
Report
Distributed LLM Inference with llm-d
An introduction to llm-d, an open-source LLM-aware router that intelligently schedules requests across inference engines like vLLM using KV cache locality and GPU utilization.
Read more
Report
Exploring Speculative Decoding: From Concept to Implementation
In this post, we explore speculative decoding through a concrete vLLM-focused implementation, covering draft models, EAGLE, MTP, and the tradeoffs involved.
Read more
Report
Exploring Mixture of Experts: From Concept to Inference Engine
In this post, we dabble in Mixture of Experts (MoE) models through a concrete nano-vLLM implementation, exploring Triton kernels, expert parallelism, and other fun things.
Read more
Report
Deep Dive into Efficient LLM Inference with nano-vLLM
A look inside a lightweight implementation of vLLM. KV cache, paged attention, tensor parallelism &multi-GPU support, etc.
Read more
Report
Pro access
Upgrade to see all 18 mentions
Upgrade to a paid plan to read every media mention of this company - funding news, awards, product launches and press releases from all the outlets writing about it.
Every media mention and press release
Funding news, awards and product launches
Fresh coverage from every outlet writing about the company
Upgrade now
Cancel anytime. Secure checkout. Instant activation.
