Are AI Labs Pelicanmaxxing?
Summary
Dylan Castillo investigates whether AI labs are pelicanmaxxing by replicating Simon Willison’s pelican-on-a-bicycle benchmark across seven frontier models, evaluating 1,008 SVG prompts. The study uses a three-stage pipeline (rendering, judging with an LLM, and feature extraction) and applies regression analysis to detect lab-level boosts. The results find little evidence of systematic pelicanmaxxing, with only a single, non-significant signal for one model in the pelican-bicycle cell and notable discussion of limitations like reliance on one judge and potential SVGmaxxing.