❯ Karpathy Announces the Pelican Test Is Retired: 1 Million Tokens to Have Opus 5 Render The Lord of the Rings as a Playable Scene
BENCHMARK ENDAndrej Karpathy shared experimental data on social media and said the era of testing large models with “draw a pelican riding a bicycle” is nearly over. He fed Opus 5 the opening passage of The Lord of the Rings with a 1 million token budget (roughly $10), asking it to do a three.js render. The model ran for about two hours and wrote 5,500 lines of code, programmatically turning the passage into a scene you can walk through in the browser. The source code is open-sourced, and you can open it and play with it directly.
WORLD SHIFTThe judgment he offers is more worth remembering than the experiment itself: models are moving from generating single artifacts to building highly customized entire worlds on demand — but they still lack the native ability to perceive and audit what they have constructed. This is a concrete engineering gap, not a philosophical reflection: whether the picture rendered from 5,500 lines of code is correct, whether anything clips through geometry, whether it is faithful to the original text — the model cannot see any of it; only a human can open a browser and look. Simon Willison — the originator of the “pelican riding a bicycle” test — also joined the discussion. The reason the old test no longer works is simple: frontier models can all draw convincingly now, and that prompt can no longer surface any difference.
EVAL COSTSWhat needs to change in step is the cost scale of evaluation. A test that costs a few hundred tokens per question and produces results in seconds cannot assess a model that can work continuously for two hours; and at $10 per run, no leaderboard can realistically be updated daily. Teams doing model evaluation now face a new constraint: who defines the scoring standard for long tasks. Human acceptance of 5,500 lines of code is unrealistic, and having another model verify it just loops back to the origin: “the model cannot see its own output.”
▪ SIGNALModels can already build worlds but have not yet grown eyes — the bottleneck in evaluation has shifted from posing the question to verifying the result.