Overview
News
Technologies
Salaries
Products
People
Growth
Financials

Overview

Security engineer writing about homelab Kubernetes, self-hosted AI, and tokens-per-second benchmarks for open-weight LLMs on real hardware.

News

News calebcoffie.com 4 months ago
Benchmarking llama.cpp's brand-new MTP support on Strix Halo
After llama.cpp merged Multi-Token Prediction (MTP) speculative decoding support, I benchmarked Qwen3.6 27B and 35B-A3B on Strix Halo and an RTX 3090. Up to 2.44 speedup, lossless output, build-from-source steps included.
Read more
Report