July 20, 2026
The Benchmarks They Don’t Bring to the Announcement
This post is part of my Medium blog.
---2"
Your Expensive Frontier Model Still Can’t Read a Clock
Every time a new frontier model ships, the announcement follows the same template. Bar exam benchmarks. Medical board benchmarks. SWE-bench. GPQA. A chart rising to the right. The model is getting smarter. Look at these scores.
But does anyone ever announce the results of a new model on ClockBench? “Good news customers, we can read a clock correctly 66% of the time.”
And every release is also accompanied by a whole series of “Us vs. Them” on various levels. Is Anthropic ahead of OpenAI? What about Google? Are open source models better than proprietary models? How far behind is country A vs. country B? … All of these news stories generate clicks, for sure, and they are all based on a series of often manipulated tests.
Celebrate the Antibenchmark (AI assist from Gemini)To be clear, I’m not questioning all benchmarks. I depend on a few of them, and it often takes a few days for the results for most benchmarks to be updated after a release. I do not trust user reports: “Oh, this new model feels better than this other model.” I don’t trust general news reporting: “Haven’t you heard that the new model is better than…”
Every model release triggers a tsunami of LinkedIn posts from people who have used a model for all of 30 minutes: “This is a revolution!”
I’m tired of it. Not because the scores are fake — they’re not. But because the tests are chosen by the people selling the model. You don’t bring a test you’re going to fail to the announcement blog post, and there are still simple tests that most of these models fail.
I’ve started calling the tests they don’t bring antibenchmarks — benchmarks designed to be trivial for humans and revealing for AI. They give me some hope that there’s still something we’re good at. And the picture they paint is more interesting than the launch decks.
Take SimpleBench. It asks questions a teenager answers without thinking. “You can’t eat cookies if they’re gone.” Last year, the best models scored 42%. Humans scored 84%. Then Fable 5 shipped in June and scored 81.9% — the first model to break 80%. Human baseline is 83.7%. The gap is under two points. I’m sure models will improve here, but you don’t see a new model announcement with a comparison to the Human baseline.
Or ClockBench — reading an analog clock. This one is my favorite because it’s such a trivial task for us simple, non-GPU-based humans, and we have a baseline score of 89%. The best model at the time of writing GPT 5.6 Sol Max at 66%. Still failing at something a six-year-old does, but the gap is narrowing. The Stanford AI Index calls this “jagged intelligence” — AI can win a gold medal at the International Mathematical Olympiad but still can’t reliably tell time.
So you can’t just say “AI can’t do simple things.” Some simple things are learned. The picture is more complicated than that.
But every time one antibenchmark gets cracked, another appears, showing the same floor-level failure. SimpleBench closed. A new one opened. And this one is brutal.
It’s called ARC-AGI-3, and it doesn’t ask the model to answer questions. It drops the model into a game it’s never seen — a grid with objects and colors, no instructions, no rules posted. The model has to figure out what it’s supposed to do by interacting. Try a move. See what happens. Build a mental model of how this world works. Plan a sequence of actions. Win.
You do this constantly. You walk into a restaurant you’ve never been to and figure out in seconds whether you order at the counter or sit down. Nobody tells you. You read the room — the layout, where other people are sitting, whether there’s a register near the door. A toddler does this with every new object they encounter. It’s the most basic thing intelligence does: encounter novelty, make sense of it, act.
Every human who takes this test solves it. Not most. All of them. 100%.
The latest frontier models? GPT Sol leads at 8% — and that’s the new state of the art for the ARC-AGI-3 competition.
This goes back to Moravec’s paradox, written in 1988: the hard problems are easy, and the easy problems are hard. Reasoning, logic, language — the things we associate with intelligence — are the shallow end. Evolutionarily, they are very recent, and they turn out to be computationally easy. Math, accounting, logic, reason: there’s a computer for that.
The deep end is perception, feeling, and exploration: walking into a new situation, figuring out what’s going on, acting. That’s millions of years of evolution. We don’t put it on benchmarks because it’s so basic we forgot it counts as intelligence.
Apple’s “Illusion of Thinking” paper showed the mechanism. Reasoning models don’t reason. They retrieve. When you make a problem slightly different from what they’ve seen — not harder, just different — they collapse, and this points to something that people understand about this emergent technology.
Models are not understanding as much as they are just aggregating and following patterns, and they are so good at this it often appears as if something has been “learned” or “understood” — and it’s important to realize that this isn’t the case. No matter how many hundreds of billions are spent on infrastructure for AI, it isn’t understanding as much as it is following.
Here’s the collective hallucination. Not that AI is dumb — it isn’t. The hallucination is that we know what “getting smarter” means. We watch a model crack the bar exam, crack SimpleBench, get better at reading clocks, and we draw a line through those data points and call it “intelligence.” But the line only goes through the things we’re measuring. We’re not measuring whether the thing underneath is comprehending. We’re measuring whether the outputs look right.
A model that passes the bar exam but can’t walk into a room and make a compelling argument to a jury. It also won’t be able to read the “analog” clock on the wall.
The gap isn’t at the frontier. It’s on the floor. Next time you read about a new thousand trillion dollar frontier model release, ask if it’s better than a 6-year-old at reading the clock.
In The Slop Codex, this is the Benchmark Theater — the field-guide monster for the gap between the numbers that get announced and the numbers that would actually matter. The counter-move: ask what the model can't do, and whether the benchmark would catch it if it couldn't.