The tech world has a new favorite word, and it means being cheap. “Tokenomics” is the art of not overpaying for AI — and this weekend’s Exchange cover made it official.
A token is the unit that measures AI usage — every request you send a model burns tokens, and every token costs money. A few months ago, big AI bills were a badge of honor. Companies rewarded employees for “tokenmaxxing” and posted leaderboards showing who spent the most. Now they are thriftmaxxing.
The Lamborghini problem
Fed up with ballooning costs, companies big and small are switching to lower-priced models, including some built in China. Many keep OpenAI and Anthropic’s products too — they just shop a la carte.
Mike Saeks, field chief technology officer at Cursor, put it plainly: the most powerful models are not necessary for mundane tasks. “It’s like driving a Lamborghini to go to the grocery store to pick up milk,” he said. His numbers are the whole argument. Building a browser from scratch on one premium model cost a little more than $10,000. Doing the same job with Cursor’s own coding model plus a rival’s cost $1,339.
“The best model for a task used to change every few months,” Saeks said. “Now it feels like it’s happening multiple times per week.”
Zero loyalty, lots of free samples
“There’s zero loyalty that I’m seeing,” says Marty Kausas, chief executive of the AI support platform Pylon. “It really feels like a bloodbath right now.” Vendors are cutting deals “like crazy” to keep customers. Kausas estimates Pylon has received about $1.6 million in free tokens from one vendor this year, $65,000 from another and $10,000 from a third.
The pattern repeats everywhere. Zoom has used Meta’s open-weight Llama model for three years and saved substantially by fine-tuning it. At the data platform Hex, roughly half of customers adopted a model from China-based Moonshot in a two-week stretch. At Brale, the CEO did the math on staying premium — “it was going to be like 100 grand per day” — and switched. The premium models still get used. They just don’t do everything anymore.
There is a geopolitical layer too. The best-known American models are closed — tightly controlled by their makers. China’s are often open-weight — downloadable and customizable. Some officials want the Chinese models banned. On Friday a group including Nvidia (NVDA) and Microsoft signed a letter urging caution and supporting open models.
Why the plumbing is the story
Here is the investing pattern underneath the buzzword. A layer of the computing stack that used to be a black box is being unbundled, priced and shopped. The winners in that kind of shift are rarely the loudest brand. They are the venues and switchboards that get paid on volume, whichever supplier is fashionable this quarter. The losers are whoever was quietly charging for the friction.
The cheap-model wave has turned the AI race on its head. It threatens the heady valuations of OpenAI and Anthropic as they prepare to go public. Their counterattack — lock-in partnerships, tens of thousands of dollars in incentives, heavily subsidized usage — is a description of a price war, not a monopoly.
Our discipline here is simple and boring. Infrastructure stories are wonderful businesses and terrible entry points, because the multiple usually pays for the narrative rather than the volume. If model-agnostic plumbing ever prints durable margins — not subsidized ones — we will revisit with numbers instead of adjectives.
