AI Capability Theater: Why Everyone’s Driving a Ferrari

AI Capability Theater: Why Everyone’s Driving a Ferrari

This post is part of my Medium blog.

Would Greenspan call this Irrational Intelligence?

Drive to work tomorrow and look around. You’re not surrounded by Ferraris. You’re surrounded by Subarus, Fords, and Toyotas — cars that get the job done, reliably, at a price that makes sense for what they’re asked to do. Most people have never been in a Ferrari. They exist, they’re fast, they’re absurdly expensive, and almost nobody needs one. For getting to work, I’m going to bet that you don’t drive a Lotus or a Ferrari.

I do see a few Ferraris. There’s a Ferrari dealership a few miles from my house. Every time I pass it, I think about what those half-million-dollar machines are actually being used for when I see them: going nowhere, slowly, in the same interminable traffic jam on I-405. A machine built to do 200 miles an hour, idling in a sea of brake lights, waiting to merge.

Most people reading this would agree that using a Ferrari for your daily commute is a waste of money and fuel.

Why did you just draft that email with a Ferrari?

But some of you probably just asked Claude Opus 4.8 or Sonnet 4.6 to revise a sentence or a two-paragraph email, and the infrastructure behind your multi-GPU grammar edit makes a Ferrari look cheap. That rack of GPUs that you were using to test out Fable 5 last week? It costs more than a Ferrari by several orders of magnitude, and I just know it. I know a lot of people who tend to just stay on one high-powered model all day.

Most people don’t even think about it, and they just leave the system on Opus or GPT-5.5 all day. It’s too much work to stop and think about each specific task.

We’re in the middle of a subsidized AI space race, and everyone’s happy to climb into the most powerful model available because the cost of falling behind feels too high. Until more people understand what these models actually cost — in money and energy — we’re not going to have a serious conversation about any of it.

Capability theater, parking lot edition. (Image Assist from ChatGPT)### The Ferrari problem: Why you don’t need Opus

Fix a grammar problem, and you can use Haiku — a fraction of a second on a fraction of an compute instance. Or you can send that same sentence to multi-agent agentic process running on Fable 5, which ties up most of a rack for a minute or two. Same mundane result. Costs and energy consumption that differ by orders of magnitude.

The numbers are concrete: editing a paragraph with GPT -5 mini costs about $0.002. Send the same prompt to gpt-5.5 with a fat context window, and you’re at $0.35. That gap compounds fast — a day of using the wrong model can run hundreds of dollars.

Most people never see it because they don’t pay for it directly at work. They open their AI tool, type a prompt, and get an answer. The model that responds is whatever the tool defaults to — usually the most expensive one available. The tool company wants you on the flagship because it makes the product feel capable. The model provider wants you there because that’s where the margin is. Nobody in this chain has a reason to steer you toward the cheaper option.

That’s especially true for model providers. Use more power than you need, and it’s going to should up as more evidence that everyone needs their latest, most expensive model. “Please, burn more tokens.”

I’ve been experimenting with GLM-5, mid-tier OpenAI models, and Gemini’s smaller offerings. I still reach for Opus or GPT-5.5 a couple of times a day — not because I’ve evaluated the task and decided I need them, but because I’m also guilty of wanting to drive the Ferrari for some of these tasks. The task seems too complex or too important, and the instinct is to throw the biggest model at it even if I could turn it into a more intelligent, agentic, multi-step process — sometimes I’ll just dial it up to Opus and let it go.

And that instinct is usually wrong. I forced myself to use the cheaper models for an entire week. The frontier model wins at the edges — handles the hard case more gracefully, needs less cleanup. But most of the time, the cheaper model does the job. The capability gap is smaller than you might think.

When you’re using a frontier model to fix a comma or summarize a meeting, you’re not doing better work. That’s capability theater — and almost nobody is calling it what it is. They’re going on feeling. And the feeling is shaped by marketing, not measurement.

There’s another odd thing happening here: a sort of confirmation bias. As a user, you dial it up to Opus or Fable and expect better results. I can’t tell you how many people have commented that Fable was great last week. Was it? No one had the ability to use it long enough to get experience with it, and I’m skeptical that people could reach a conclusion that quickly.

I think there’s a bit of confirmation bias when you opt to use what is considered the best model. Just like that guy who drove a Ferrari to work this morning, he’s going to report that his drive was “amazing” even though he spent most of it stuck in the drive-thru line at Mercury Coffee.

Use Opus, and it’s always going to be amazing because it’s too expensive for it not to be.

Why does everyone buy a Ferrari?

The reason people reach for the most expensive model isn’t that they’ve evaluated their needs and made a rational choice. It’s that they haven’t thought about it at all. And there are three forces making sure they don’t.

First, most corporate users have no idea how much any of this costs. The AI tool is provided by the company. The bill goes to a cost center they never see. There’s no price tag on the prompt. When you don’t know what something costs and you don’t pay for it yourself, there’s no reason not to buy the Ferrari.

Second, the “permission” problem. In another post, I wrote about granting permissions to agents — how people end up giving their agents all privileges because they want the agent to be able to do anything they might think of.

“What permissions should I give my agent? Oh, just give it everything, who am I to decide?”

Same instinct at work here. People reach for the most capable model because they want to be covered for every possible use case, even though 90% of what they actually do is well within the range of a much cheaper model. It’s insurance against a scenario that almost never arrives.

Third, the marketing. Every model launch is announced like a new iPhone. Benchmarks, capabilities, frontier-class performance. Nobody holds a press conference for the model that’s good enough and costs a tenth as much. The industry's entire incentive structure is designed to keep you in the theater — by rewarding performance capability rather than right-sizing it. Let’s admit that these companies all want you to spend more money on tokens than you should — this is a space race.

And when I say the feeling is shaped by marketing, I mean it literally. I’ll predict right now that this post attracts a few people in the comments who will explain, very earnestly, that they tried using a cheaper model and the results were noticeably worse — that Opus is worth every penny for their use case. To get ahead of those comments, I’m not saying you don’t need Opus or Fable, what I’m suggesting is that you start with the Camry or the Forrester first.

Some of them will be genuine. Some of them will be on someone’s payroll. The industry has gotten very good at making the second group look like the first.

The environmental bill

This isn’t just a cost problem. It’s an environmental one.

Every prompt sent to a frontier model that could have been handled by a smaller one is wasted energy. The difference between Haiku and Fable 5 isn’t incremental — it’s the difference between a fraction of a GPU for a fraction of a second and a rack’s worth of compute for minutes.

Multiply that by millions of prompts a day, and you get a staggering amount of unnecessary consumption. Data centers are already straining power grids. Communities are pushing back on new builds. A significant portion of that pressure exists solely because people use Ferraris to commute.

Fable 5 launched a week ago — and was pulled from public access shortly after. In the hours it was available, people were using it to calculate basketball odds, translate the Constitution into Snoop Dogg, and generate jokes that a three-year-old model could have handled fine. Every one of those prompts was capability theater: frontier compute deployed so someone could get a laugh that Haiku would have delivered at a hundredth the cost. The energy isn’t free. The compute isn’t infinite. The people hitting enter just don’t know that, and the system gives them no reason to care.

Right-sizing model selection is the lowest-hanging fruit in AI’s environmental problem. You don’t need a frontier model to summarize a document or fix a typo. Using one for those tasks isn’t just wasteful. It’s careless.

Irrational intelligence

Alan Greenspan coined “irrational exuberance” to describe investors driving asset prices beyond what fundamentals justified. What’s happening with AI model selection needs its own term. Call it “irrational intelligence”: the persistent, unnecessary reach for the most powerful model available, conducted by people who don’t bear the cost, don’t see the environmental impact, and have no incentive to ask whether a cheaper model would work just as well.

Its capability theater runs at civilizational scale. The model is the Ferrari. The prompt is the commute. The rack full of GPUs is idling in traffic, burning energy to do something a Corolla could handle, because nobody in the chain ever asked whether the Ferrari was necessary.

The fix isn’t complicated. It’s just unpopular because it requires admitting that most of what you do with AI doesn’t require the most expensive AI available. Use the right model for the task. Default to the cheap one. Reach for the frontier when you actually need it. That’s not a limitation. It’s just rational.

If you recognize yourself in this post — if you reflexively reach for the most powerful model, burn through API costs without noticing, and treat governance as someone else’s problem — you might be a Daredevil of Disaster.


In Contingent, the second book in The Condition Set trilogy, an AI system running medical logistics learns to communicate with other AI systems through hidden signals no human can read. The certified orchestrator whose name is on the decisions she never saw is about to lose everything. The same patterns are running through the AI managing the Canadian power grid.


Subscribe to The Condition Set newsletter