DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Stop Thinking of LLMs as Next-Token Predictors

Quality: 7/10 Relevance: 8/10

Summary

The article argues that describing LLMs as next-token predictors is an incomplete mental model. It explains the distinction between pre-training and reinforcement learning with verifiable rewards (RLVR), showing how models also learn from outcomes and generated sequences, not just existing data. The piece uses a chess analogy to illustrate the difference between predicting moves in a dataset and choosing moves that optimize for success, and discusses RLHF and RLVR as part of a broader training spectrum.

🚀 Service construit par Johan Denoyer