Why it matters: Many people report half this throughput on the same hardware; the specific quant and serving settings are the difference between a sluggish and a genuinely usable local coding assistant.
How to apply: Match the poster's quant (RedHatAI build) and vLLM version/flags before assuming you need more or better GPUs for this model.