Exploring Claude/GPT Knowledge Cutoffs and Pre-Training Timelines
Summary
Shrivu Shankar analyzes how frontier models are trained in three stages—pre-training on broad data, domain-specific fine-tuning, and post-training to craft the assistant persona—and describes methods to infer training timelines and knowledge cutoffs through probing, data mixtures, and self-identification prompts. The piece argues that probing can reveal approximate model histories and compares patterns across model families, though it emphasizes these inferences are approximate due to limited public ground truth.