Why it matters: Expert cache misses, not raw compute, are the hidden tax on consumer MoE inference; prefetching is a cheap software win on hardware you already own.
How to apply: Clone github.com/Niko1221/Strata, deliberately shrink your expert cache to measure your miss cost, then enable prefetch and re-measure tokens/s.