The AI Price War Is Rewriting Your Software Budget

Akshay Bhimani
Akshay Bhimani
The AI Price War Is Rewriting Your Software Budget

OpenAI released GPT-5.6 on 9 July with an entry tier priced at $1 per million input tokens, nine days after Anthropic launched Claude Sonnet 5 at an introductory $2 input and $10 output (Developers Digest, July 2026). Frontier-grade AI now costs a fraction of what GPT-4 charged at its 2023 debut. For small and mid-sized businesses, this summer’s price war resets the cost of the AI features and automation across your software stack.

200x per year

median fall in the price of fixed-capability LLM inference since January 2024 (Epoch AI 2026)

$1 vs $30

GPT-5.6 Luna’s input price per million tokens today against GPT-4’s at launch in March 2023 (Developers Digest 2026; Introl 2026)

61%

organisations that cut projects in the past 12 months after unplanned SaaS cost increases (Zylo 2026)

What Is Actually Happening Right Now


Three price moves landed within six weeks. OpenAI’s GPT-5.6 ships in three tiers: Sol at $5 input and $30 output per million tokens, Terra at $2.50 and $15, and Luna at $1 and $6 (AI Pricing Guru, July 2026). Anthropic’s Claude Sonnet 5, launched on 30 June, holds introductory pricing of $2 and $10 until 31 August before it settles at $3 and $15 (Developers Digest, July 2026). DeepSeek V4 Flash undercuts them all at $0.14 and $0.28.

The pattern matters more than any single number. Every major provider now sells a budget tier capable of work that needed a flagship model a year ago. That is competition, not generosity, and it hands buyers bargaining power they did not have at their last renewal.

Why SME Software Budgets Are Most Affected


Here is the uncomfortable part: raw model prices are falling, yet most SME software bills are still going up.

  • AI uplifts of 20–37% now appear on renewal quotes even as vendors’ own model costs drop (Lynton 2026)
  • SaaS spend rose nearly 8% in a year as vendors monetised AI features (Zylo 2026)
  • Hybrid subscription-plus-usage pricing shifts overage risk onto the buyer, often mid-contract
  • Custom AI quotes written in 2024 assume token costs that no longer exist

What Changed Between 2024 and 2026


The speed of the collapse is the story. Epoch AI’s analysis of six benchmarks found the price of fixed-capability inference fell between 9 and 900 times per year, and the median rate since January 2024 is 200 times per year – far steeper than the roughly 10x annual decline Andreessen Horowitz measured when it coined the term “LLMflation” in late 2024 (Epoch AI 2026; a16z 2024).

The comparison anchor: GPT-4 launched in March 2023 at $30 per million input tokens and $60 for output (Introl 2026). GPT-5.6 Luna, a stronger model, now costs $1 and $6. One caveat for buyers. Developers Digest notes that Claude models from Opus 4.7 onward use a tokeniser that can produce up to 35% more tokens for the same text, so compare the cost of a finished task, not the headline rate (Developers Digest, July 2026).

Regulation moved too. The EU Artificial Intelligence Act (EU, in force since August 2024) began enforcing obligations on general-purpose AI model providers on 2 August 2025, and most remaining obligations take effect on 2 August 2026 (European Commission). Compliance overhead is one reason vendor list prices have not fallen as fast as raw tokens.

Still paying 2024 prices for AI that now costs pennies?


Most renewal quotes on your desk were priced before this year’s cuts. HMMBiz reviews what your tools actually spend on AI and flags where a custom build now beats the vendor add-on.

Talk to Our AI & Automation Team

What Your Business Should Do Right Now


  • Ask every SaaS vendor at renewal which model powers its AI features and when that price was last reviewed. A vague answer is your opening to negotiate.
  • Re-quote any chatbot or automation project scoped before 2025. The token costs behind the old estimate have fallen by an order of magnitude.
  • Route routine, high-volume tasks to budget tiers such as GPT-5.6 Luna or DeepSeek V4 Flash. Reserve flagship models for complex reasoning.
  • Set usage caps and alerts before rolling out any hybrid-priced tool. Overage charges, not subscriptions, cause most surprise costs.
  • Benchmark one real workload across two providers every quarter. Loyalty at stale prices is the most expensive habit in AI buying.

HMMBiz Perspective


HMMBiz builds chatbots and workflow automation for SMEs across five markets, and the most common client question has flipped this year: no longer “can we afford AI” but “why is our software bill still rising while model prices fall”. The gap between falling input costs and flat retail pricing is where most SME budgets leak. Vendors reprice slowly, and buyers rarely push back.

HMMBiz runs Chato, its own AI website chatbot, on the same commodity tokens its clients buy, which keeps project pricing tied to current rates rather than last year’s assumptions.

Make the price war work for you

HMMBiz has been building AI chatbots and automations since tokens cost real money. Now that they barely do, the project you shelved in 2024 deserves a second quote.


Get a Fresh Quote

India · USA · UK · Australia · UAE

Frequently Asked Questions


Why are AI model prices falling so fast?

Competition and efficiency gains. Epoch AI found the price of fixed-capability inference fell between 9 and 900 times per year across benchmarks, with the median at 200 times per year since January 2024 (Epoch AI 2026). Cheaper inference hardware, smaller specialised models and low-cost entrants such as DeepSeek all push in the same direction. Falling rates do not guarantee falling bills, because usage tends to grow faster than prices drop.

Does the EU AI Act change how my business buys AI features?

Yes, if you operate in or sell into the EU. The EU Artificial Intelligence Act (EU, in force since August 2024) applied obligations to general-purpose AI model providers from 2 August 2025, with most remaining obligations effective 2 August 2026, and penalties of up to €35 million or 7% of global turnover for the most serious breaches. Ask vendors which model versions power their features and where your data is processed; HMMBiz includes both questions in its standard vendor review.

How quickly can HMMBiz build a custom AI chatbot at current prices?

A scoped website chatbot typically goes from kick-off to launch in two to four weeks. HMMBiz starts with a one-week discovery covering use cases, data sources and model selection, then quotes against current per-token rates rather than a fixed 2024-era estimate. Clients get the working bot, an admin dashboard and a monthly usage report so running costs stay visible.

What is a token, and why is AI priced per million of them?

A token is a chunk of text, roughly three-quarters of an English word, that a model reads or writes. Providers charge separately for input tokens (what you send) and output tokens (what comes back), quoted per million. Two models with similar list prices can produce very different bills because their tokenisers split text differently, so always compare the cost of a finished task rather than the per-token rate.

REFERENCES

1. Developers Digest – “Frontier Model API Pricing, July 2026: Claude vs OpenAI vs Gemini vs DeepSeek” (updated 16 July 2026)

2. AI Pricing Guru – “GPT-5.6 Pricing (July 2026): Sol $5, Terra $2.50, Luna $1 per 1M” (updated 20 July 2026)

3. Epoch AI – “LLM inference prices have fallen rapidly but unequally across tasks

4. Introl – “Inference Unit Economics: The True Cost Per Million Tokens” (documents GPT-4’s March 2023 launch pricing of $30/$60 per million tokens)

5. Zylo – “2026 SaaS Pricing Trends Driving Up Enterprise Costs

6. Lynton – “The 2026 SaaS Pricing Squeeze: When Your Vendor’s Commodity Problem Becomes Your Budget Problem

7. Andreessen Horowitz – “Welcome to LLMflation: LLM inference cost is going down fast” (November 2024)

8. European Commission – “AI Act: Regulatory framework for AI” (application timeline)

Recent News


LinkedIn began testing Collaborative Posts at Cannes Lions in June, a format that lets two or more accounts co-author a […]

Google now answers a majority of searches on the results page itself, and the click that used to follow often […]

Scroll to Top