Kimi K3, and what we can still learn from the pelican benchmark
Summary
Simon Willison analyzes Moonshot AI's Kimi K3 (2.8T parameters) and the pelican benchmark, noting K3's pricing and open-weight promises. He argues that while the pelican test offers insights into prompting cost and model behavior, it no longer reliably predicts real-world agentic capabilities, especially for tool-calling and longer interactions, though it remains useful as a hands-on exploration.