Picture a fast typist working next to a careful editor. The small model rattles off a guess at what comes next, and the bigger model checks a whole batch of those guesses at once instead of writing every word itself. When the guesses are good, and they usually are, you get the quality of the big model at close to the speed of the small one. The technique has a fancy name, speculative decoding, but that is the whole intuition: guess ahead, verify in bulk.
On a plain 8GB laptop with no dedicated graphics card, that gets you the first word back in about seven tenths of a second, and roughly twelve words a second after that. Drop in even a modest GPU and it climbs to about thirty-five. Fast enough that you stop thinking about the machine and just talk to it.
The full write-up gets into the quantization choices and why we did not have to trade away answer quality to hit those numbers. Running well on 8GB turned out to be a math problem, not a compromise.
This update lives on Nunba. Join the community to be part of the next one.