Big Picture
Alright, happy Friday! I think it’s time to bring the AI discussion back to the more technical level. Less about Macro this week. It is getting old. So on AI it is where are we, and how are the models actually making money... or not? Anthropic has already confidentially filed to go public, private AI companies keep getting funded at massive valuations, and the economics are starting to matter a lot more. So I think we need a reset to gauge current conditions in a world where AI is dominating much of the oxygen.
So to start here, most of the AI debate still spends its time on what the models can do. Whose model is better? What can one do that another can’t? The wow moments that drive demand. In many ways using scare tactics (Dario and Sam). “Our model is so good we cannot release it”.
If we look just 9 months ago Anthropic and Claude came out swinging, really pushing usage from chat and search toward more agentic workflows. OpenClaw, Cowork, and other similar-ish offerings...
Since then we have Muse and Instinct, essentially always-on personal agents that can actually go do things for you. The distinction today is that these are mostly single-player products. They work for you, not really with your family or your team. It is actually a pretty big technical, orchestration and governance challenge.
Companies we know that have created tools to assist here are Block with Buzz which is open source, and Salesforce with Slack trying to build the shared surface where humans and agents work together. All of this means more workloads, more inference, more tokens. That is obvious.
But less attention has been paid to the cost side of the equation. That is starting to change. The headlines are slowly moving from “what can the model do?” to “how efficiently can the model do it?” That matters because the labs now have to balance three things at once: growth, frontier intelligence, and some path toward economics that actually work. Kind of reminds me of Uber / Lyft days, or WeWork vs its field. Whoever has more money + execution obviously has the chance to become the dominant player.
And that is where things get interesting, at least for me. A lot of these model companies are stuck between being the intelligence layer and being the application. But in just the last month Muse is one of the clearest examples of the application starting to abstract the model away. You are not choosing a model for every task. You are just trying to get something done at a price you can afford. This is how the world has worked forever.
Think about solar panels on your roof while you’re still connected to the grid. Once it’s set up, you are not constantly deciding whether the house should run on solar or grid power. You just want reliable energy at the best price. AI is starting to feel the same way. The model becomes a backend layer and the product is the experience on top. This happened during internet boom, cloud boom, mobile boom. Infra comes forward then takes back seat to everything else.
So this week we are walking the full cost stack of an AI model. Where the build money goes, what each layer costs to run, what happens to demand when price collapses, and where the value ends up when the model layer gets commoditized. We’ll give you our view at the end but the short version is this: we think cost at the model layer continues to squeeze lower. Models become the internet a smart yet commoditized layer to do things. And the value migrates to whoever owns distribution, proprietary context, orchestration, and inference scale.
First, the three places every model dollar goes. You can see that on the first chart.
[1] Every AI dollar goes to one of three places.
At a high level, there are three places compute goes. And yes some readers here like hedge fund managers, and financial analysts or even some of the technology experts know this. But i do think its important to refresh your mental model and ensure the logic from 9 months ago still holds today. To be frank, it’s becoming ‘does logic from 4 weeks ago hold today’ this is moving that fast… so lets go.
Training is where the base model learns broad patterns from huge datasets.
Post-training is where that raw model gets shaped into something useful, through fine-tuning, reinforcement learning, tool use, and a lot of trial and error.
And inference is the model actually doing the work for the user, one request at a time, so it becomes the recurring variable cost that scales with usage.
Plain English: training is med school. Post-training is residency. Inference is every patient after that. Couldn’t think of a better analogy lol.
The lines are getting blurrier because modern post-training can involve the model generating and scoring enormous numbers of its own attempts. Operationally, that can start to look a lot like inference even though it is still part of training.
[2] The famous training run is a small line item.
Now the number everyone quotes is usually the final training run. But Epoch AI’s work suggests that is only a fraction of total R&D compute spend.
Their estimates put OpenAI at 9.6%, Z.ai at 12.3%, and MiniMax at 22.6%. The rest is everything around the run: experiments, synthetic data, failed attempts, post-training, research, and all the work required to figure out what is worth training in the first place.
Think of it like a movie made in Hollywood, or that may be a bad example when we review this piece in 10 years..but the final shoot is expensive, but it is not the whole production budget. That is key the next time someone quotes one giant training-run number as the total cost of building the model. So again keep that in mind and the image below shows this…
[3] And the build bill keeps growing.
Indexed to 2020, the trend for frontier language-model training cost goes from 1 to roughly 1,838 by 2026.
That is about 3.5x per year.
So yes, efficiency is improving everywhere, but the labs are still spending more to push the frontier. The arms race has not gotten cheaper. Building is getting more expensive. It is almost like CAC, cost of acquiring a customer… but that customer is spending less per unit of value you provide…
Now lets flip to the other side of the ledger here, because this is probably the most important chart in the piece. I think at least.
[4] The cost of fixed intelligence is collapsing.
Take a fixed level of intelligence that cost $100 to use in late 2022.
Now hold the capability constant and track the cheapest model that can deliver that same performance over time. Estimates suggest that same level of intelligence costs roughly $0.007 today. That is a 47% decline in cost per quarter, or roughly 13x cheaper every year.
Yes, you read that right.
This is the crazy tension sitting at the center of AI economics right now. The labs are spending exponentially more to push the frontier forward, while last year’s level of intelligence is becoming dramatically cheaper to use. It is wild really.
So both things are happening at once: the level of intelligence is rising, while the cost of existing intelligence is collapsing. That is why the economics of AI can look completely different depending on which side you are staring at. But we know for sure that we do not pay estimates. We pay invoices for this stuff. So let’s look at what businesses are actually spending…
[5] The proof is in real spend data.
So we turn to some latest data. Ramp AI Index now tracks token volume, token spend, and effective token pricing across a sample of businesses using its Token Spend Management product. It is not the entire Ramp customer base, but it gives us a useful real-world read on the direction. Conceptually and speaking to various public and private companies this direction at least lines up.
Since January, token volume in that sample is up roughly 10x. Token spend is up too, but nowhere near as fast. That is the point. Companies are consuming much more intelligence without the bill scaling one-for-one with usage.
That is deflation in actual TOTAL business spend. The next two charts help explain why I think.
[6] Not all tokens cost the same.
So reminder, a token is a small chunk of text, roughly a word fragment, and it is the unit AI is generally priced in. But a token is not just a token.
GPT-5.6 Sol lists cached input at $0.40 per million tokens, new input at $4.00, and output at $20.00. Same model, same moment, a 50x spread between cached input and output.
Cached input is context the model has already processed and can reuse. This matters a lot for agents because they often revisit the same instructions, files, and conversation history over and over. Caching does not make every agentic workload cheaper, since agents can also generate a lot more output and reasoning tokens, but it is one big reason effective token prices can fall as usage scales.
[7] And you choose how long the model thinks.
Reasoning settings are the second lever. Move GPT-5.6 Sol from LOW to MAX and Artificial Analysis has its intelligence score going from 34 to 47, while cost per task goes from $0.26 to $1.99. That is 7.6x the cost for roughly 38% more intelligence on their index.
You do not need MAX reasoning to answer what time a restaurant closes. The better AI products increasingly route the easy work to cheaper models and settings, then spend more compute only where it actually matters.
That is why routing is increasingly becoming a margin line. Cost discipline is becoming part of the product. Also routing is becoming a business, Stripe acquired OpenRouter, and NVIDIA acquired Hugging Face..
[8] Cheaper intelligence means more compute, not less. But still, where is the real value?
Now this helps the infrastructure debate. If intelligence gets dramatically cheaper, demand does not have to fall. It can explode.
So far estimates around AI are that the total computing power of the installed AI-chip stock is growing about 3.4x per year. Separate token-demand proxies from providers and routing platforms have been growing much faster, often on the order of 7x to 30x, with around 10x used as a rough market-wide midpoint. Those demand estimates are noisy for sure, so I would not pretend the exact gap is known.
But the setup is pretty obvious here: efficiency gains can get absorbed by more usage. That helps explain why better models and lower token prices have not automatically translated into less infrastructure demand.
It also does not mean every GPU or data-center dollar will earn a great return. It just means the bear case cannot stop at “models are getting more efficient, so we will need less compute.”
[9] Where it goes: agents doing the buying.
Now take that cheap intelligence and let it transact. We see estimates that if end-to-end agentic commerce scales, which I think it will, U.S. e-commerce could approach $2 trillion by 2030, roughly 11.5% above its baseline forecast. That does not mean $2 trillion of purchases will be made fully autonomously by agents. It means agents could become a meaningful part of discovery, comparison, and checkout across a very large transaction base. We have a view here on the opportunity for this, but that’s for another piece maybe.
Once intelligence gets cheap enough, it stops being a product you visit and starts showing up inside the transaction itself. That is when the economics move beyond model subscriptions and API calls.
[10] The landscape: who builds the model, who builds the product.
The market is splitting into two camps. Model creators build the brain. Agent platforms build the product on top. Below is the landscape as of today, public and private, across both sides. Some are playing both sides, eventually I think all will NEED to play both sides.
Net Net
Training bills keep growing while the price of fixed intelligence keeps collapsing. That sounds contradictory, but it is really the core of the whole thing here. The frontier gets more expensive to build while intelligence gets cheaper to distribute, and lower prices can create more usage.
So I do think the model layer faces real deflationary pressure, both from competitors and from each company’s own next release. But I would not jump from that to saying the model companies cannot make money. They can still capture value through distribution, product ownership, enterprise integration, proprietary context, and scale. Actually I think they will NEED TO to be around a decade from now. The question is how much of that value stays with the model itself versus the products and workflows built around it.
If this trend keeps going, the model layer starts to look more like infrastructure: essential, everywhere, and under constant price pressure. That would push more of the economics toward the layers that own the customer, the context, the workflow, and the transaction. This is why what’s happening with Muse is so important to pay close attention to.
So it is not just who has the smartest model, but who can turn cheaper intelligence into durable economics.
That’s all for this week!
About Avory & Co.
Investing Forward.
Avory specializes in high-conviction equity strategies, emphasizing Secular Growth and Transformation Stories driven by exceptional teams. Data guides decisions. We cater to high net worth investors, family offices, and institutional investors. Note: This information doesn’t constitute a recommendation to buy or sell any mentioned securities. Avory is based in Miami, Florida with clients all across the globe.
Speak to us:
Send us an email: Team@avoryco.com
Want to invest? We are on most platforms.
Want More
🎥 Avory YouTube: Channel
🎙️ Avory: Podcast
Disclaimer: Not a recommendation to purchase or sell any securities mentioned. This is for educational purposes only.












