The Inference Wall
The Inference Wall
An 8.6 GB model that serves only 7 requests a second, and the trace that says why
A primer: how an LLM actually serves a request
Benchmarks 01 — hit the wall
Experiment 01 — hit the wall
Notebooks
Shared scripts
Notebooks
Notebooks
Analysis notebooks, published alongside the posts they support.