DigiNews

Tech Watch by Johan Denoyer

← Back to articles

Kimi K3, and what we can still learn from the pelican benchmark

Quality: 8/10 Relevance: 9/10

Summary

Simon Willison analyzes Moonshot AI's Kimi K3 (2.8T parameters) and the pelican benchmark, noting K3's pricing and open-weight promises. He argues that while the pelican test offers insights into prompting cost and model behavior, it no longer reliably predicts real-world agentic capabilities, especially for tool-calling and longer interactions, though it remains useful as a hands-on exploration.

🚀 Service construit par Johan Denoyer