Gradient blur
Serverzaal met lichtstrepen die rekenkracht en snelheid uitbeelden
Blog
Artificial Intelligence
Back to overview

AI prices are falling, but for how long? How to build an architecture that absorbs a price shock

AI model prices are dropping fast, and for anyone building with AI today, that is good news. Behind those discounts, though, is a market that still runs largely on investor money.
01 - 10 - 2026

Peter Verrykt, Business Director Professional Services at Xylos, lays out the numbers and explains how to build your AI architecture so it absorbs a price shock, whichever way the market moves.

At the end of September, Anthropic and OpenAI cut their prices on the same day. Anthropic released Claude Opus 5.5, 20% cheaper than its predecessor. A few hours later, OpenAI set the price of GPT-6 Sol and Luna at half of what the previous generation cost. For anyone buying or using AI today, that is good news. But look beyond the next invoice and you see a market that sells at a loss.

A business model that runs on investors

Let's start with the numbers. OpenAI generated 5.7 billion dollars in revenue in the first quarter of 2026 and burned through 3.7 billion dollars in cash in those same three months (The Information). According to an investor presentation the Financial Times saw in mid-September, OpenAI expects a negative free cash flow of 278 billion dollars between 2026 and 2030. By 2030, some 856 billion dollars will go to compute and infrastructure. Despite a 122 billion funding round in March, the money could run out as early as 2028. That is why OpenAI is raising capital again, at a valuation of more than 1,200 billion dollars.

Anthropic does somewhat better on paper, but follows the same pattern. In February it raised 30 billion dollars at a valuation of 380 billion. Three months later came a 65 billion round at 965 billion. According to its leaked IPO prospectus, Anthropic posted an operating loss of more than 8 billion dollars last year on revenue of 4.6 billion. It has also committed to 518 billion dollars in infrastructure. Anthropic does not expect a positive free cash flow until 2028. Compute costs are falling, though: in the first quarter they still came to 71 cents per dollar of revenue, in the second quarter 56 cents. On an adjusted basis, Anthropic even made an operating profit in the second quarter. But that trajectory, too, depends on a market that keeps investing.

So the whole sector runs on the expectation that revenue will one day catch up with investment. As long as that expectation lives, the money keeps flowing. It is a classic peak in the hype cycle. A product that is sold below cost and kept afloat by capital is, of course, not sustainable in the long run.

A price war blocks the way out

Normally you get out of a situation like this by passing on the real cost price. But prices are moving the other way, and there is a clear reason for that.

Chinese models have largely closed the gap and cost only a fraction. DeepSeek V4 Flash costs 0.15 dollars per million input tokens off-peak. Alibaba's Qwen 3.5 Flash costs 0.10 dollars. By comparison, Opus 5.5 still costs 4 dollars per million input tokens after the price cut. In a comparison cited by Capacity in August, the same task cost 544 dollars on Zhipu's GLM and 4,811 dollars on Claude. That is almost nine times as much. According to Silicon Data's price index, analysed by Jefferies, the average price per million tokens fell from 2.04 to 1.16 dollars between late May and early August.

The western labs are responding with price cuts. In August, GPT-5.6 Luna dropped by 80% and Opus 5 launched at about half the price of its predecessor. At the end of September came the next round. OpenAI explicitly states that the new prices are permanent. Sanchit Vir Gogia of Greyhound Research calls it "a land grab for the default route through which enterprises buy intelligence". Gartner says models are becoming a commodity.

Whoever buys market share at a loss is counting on a later moment when they can set the price. That moment comes through consolidation, through investors losing patience or through an IPO that has to deliver returns. Nobody knows exactly how that will play out. What is certain is that today's prices are not the prices of three years from now.

Chart of the falling price per million tokens and the prices of Opus, DeepSeek and Qwen

What matters now

For companies using AI today, only one thing counts: will what you are building now still work and still pay off when the market turns? That turn can go two ways. Prices can rise when the subsidy stops, but they can also fall further, and then you pay too much if you are locked in to one vendor. On the other hand, models can disappear quickly. Opus 5 was replaced after just two months. If you build your application around one specific model, you are building on sand.

An architecture that absorbs this has a few fixed characteristics:

Six characteristics of an AI architecture that absorbs a price shock

First, there is an abstraction layer between your applications and the models. That can be an AI gateway that receives every call and forwards it to the model you have chosen. Within the Microsoft ecosystem, you can do this with the AI gateway capabilities of Azure API Management, for example, while Azure AI Foundry brings together models from different vendors. Switching models then becomes a configuration choice, without having to rebuild everything. Use vendor-specific features deliberately. Document them, so you know what to replace when the time comes.

Second, you route per task. You only need a top model for part of your requests. A classification or a summary runs fine on a model that is twenty times cheaper. You reserve the top model for the work that makes the difference. That is often where the biggest savings are, bigger than any price cut delivers.

Third, you have your own evaluation set. A set of representative tasks with expected outcomes lets you decide within a day whether a new or cheaper model is good enough. Without such a set, every model switch remains a gamble, and then you do not switch.

Fourth, your knowledge stays yours: your prompts, your context, your documentation and your data. They have to be separate from the model, so they move along to the next one.

Fifth, you keep a warning light on your running costs. Measure your costs per application, per customer and per completed task, and not just per token. Set budgets and thresholds, so you spot a runaway process or an agent stuck in a loop after an hour, and not only on the monthly invoice. What you do not measure, you cannot steer, and with consumption-based pricing the invoice quietly rises along with it.

Finally, you have an escape route. An open-weight model that you host yourself or place with a European partner covers part of your applications and serves as insurance against a price shock or a vendor changing its terms. It also helps for data that is not allowed to leave Europe.

In fact, this is all familiar ground. It is the same discipline we learned with the cloud twenty years ago: do not become dependent on one party, measure what you consume and make sure you can leave. With AI it just goes faster, and prices shift a little harder.

Use the discount to get your architecture right

With Anthropic's price cut, you pay 20% less for Opus, and even 60% less for cached input. For many organisations, that immediately frees up budget. You can spend that money on more tokens. You can also invest it in something that lasts longer than current prices: a review of your AI architecture, a gateway, cost monitoring and an evaluation set that lets you switch models tomorrow.

The price war makes experimenting cheap today. Those who use this period to get their foundations in order will be ready when the correction comes.

Xylos reviews your AI architecture with you

Want to know how agile your current AI applications are? Our AI architects look with you at where you are tied to one model or vendor, where your costs are quietly rising and which steps to take first.

Get in touch for an exploratory conversation.

About the author

Peter Verrykt is Business Director Professional Services at Xylos. He helps organisations turn data into concrete business value. He helps companies look beyond technical implementations and uses data and AI as a foundation for better decisions, greater agility and sustainable growth.