I'm (mostly) picking models on speed now, not intelligence
Summary
The article argues that speed, not raw intelligence, is becoming the primary factor in practical AI model use. It discusses throughput thresholds (around 100 tok/s) and how faster models trade some reasoning time for overall responsiveness, with examples of GLM5.2, DeepSeek, and price competition shaping the market.