Writing · Tag
1 post tagged #model-sizing.
One local LLM speed number hides the work behind it. My RTX 5090 sweep shows why short generation, long context, and code need separate rates.
Real costs, real tools, no fluff. M-F when I ship, publish, or learn something worth sending.