OpinionAi CostsUnit Economics

Prices Fell. Your Bill Didn't.

Why are AI costs rising as models get cheaper and more capable? Because better economics invite more features, users, agents, and work.

SuperPenguin Team4 min read
Prices Fell. Your Bill Didn't.

AI models got cheaper and better. Spending kept rising.

That is not a contradiction. Better models made more work worth doing. Products added features, prompts grew, more people used them, and agents began making several calls for one task.

A bill is price times volume. The AI industry watched the first number fall and underestimated how quickly the second would grow.

That is why asking “when will tokens get cheaper?” is no longer a cost strategy. A cheaper unit can create a bigger bill.

Price gets the headline. Volume gets the invoice.

Suppose a model price falls by half. If your usage stays flat, the saving is obvious. But teams do not hold usage flat.

A support group expands an assistant from agents to every customer. A coding tool moves from autocomplete to long-running agents. A product adds document search, then attaches more documents. A reliable workflow invites more users. Each improvement is reasonable on its own. Together, volume grows by five times while price falls by two.

The total bill rises by 150 percent.

Even when listed rates fall again, usage has no natural ceiling. Agents can keep working after a person would have stopped.

Snowflake showed what Jevons paradox looks like in software

In the nineteenth century, economist William Stanley Jevons observed that making coal more efficient did not necessarily reduce total coal consumption. Better economics made coal useful for more activity, so demand could grow faster than the savings.

Snowflake faced the same apparent tension in cloud data. If the company made each unit of work more efficient, would it lose consumption revenue?

In May 2021, Snowflake estimated what the tradeoff would cost. CFO Mike Scarpelli said new storage compression was expected to reduce full-year revenue by roughly $13 million because customers would need less storage. He also predicted that the better economics would lead customers to store more data and eventually drive more consumption.

By March 2022, Snowflake said it had seen this pattern before. Scarpelli said customers had historically moved more workloads after price-performance improvements, usually with a delay, and estimated that additional workloads would offset part of the latest revenue impact. Then-CEO Frank Slootman called the strategy “not philanthropy” and said the efficiency gains stimulated demand, although not immediately.

A faster, cheaper model makes a previously silly feature economical. Developers build it. Customers use it. The lower unit price creates new volume.

AI has the same shape. Cheaper inference does not only reduce the cost of an existing prompt. It changes what people attempt:

  • applications add AI to lower-value interactions
  • prompts include more context because they can
  • agents explore several paths in parallel
  • teams run evaluation and generation more often
  • users ask for work they previously did themselves
  • products serve customers who were uneconomical before

This is good news for useful AI products. It is bad news for a budget built on the assumption that next quarter's model will automatically rescue it.

Call it cheaper units, bigger bills.

Rate shopping still matters, just less than you think

Rates still matter. Negotiated discounts, caching, batch processing, and a suitable smaller model can produce real savings. Just do not bank the saving forever while behavior changes around it.

A team celebrates a 40 percent rate reduction and raises limits. Product adds two new agent steps. Customers send twice as many requests. Six weeks later, spend is higher and no one can explain why because the model comparison stopped at price per million tokens.

The relationship is simpler than the invoice:

AI spend = cost per task × number of tasks run

Cheaper tokens can lower the first number. Better models and new features increase the second. If the number of tasks grows faster than the cost of each task falls, spending rises.

Model performance changes the cost per task. A more expensive model can still cost less if it succeeds in one call instead of three. A cheap model can cost more if the workflow retries it or sends the result to a stronger model anyway. Price per million tokens matters, but it does not tell you what the task costs.

Cheaper models will not save an unmeasured product

The amount of useful work available at a given price has risen sharply. That will enable products that could not exist before. It will also tempt teams to put inference into more places and let agents do more work. Both can be true at once: each dollar buys more, and the company spends more dollars.

Snowflake understood that efficiency can expand consumption. AI buyers are now living on the other side of that equation. Waiting for the next price cut will not explain the bill if usage keeps expanding faster.

The important question is not whether token prices fell. It is whether the extra work made possible by those lower prices grew even faster. A rising bill may reflect a useful product finding demand, or a workflow quietly becoming more expensive. Either way, price alone cannot tell you which one happened.

Sources

  1. Snowflake Q1 FY2022 corrected earnings-call transcript, May 26, 2021.
  2. Snowflake Q4 FY2022 corrected earnings-call transcript, March 2, 2022.

Keep reading