The Token Price Collapse (And Why AI Costs Still Increase)
This week has showcased the continuous battle of competitors driving down the price of tokens.
First, OpenAI announced cost reduction to models GPT-5.6 Luna and GPT-5.6 Terra.
The same day, DeepSeek dropped V4-Flash – an open model that competes on price too, and free to self-host.
The chart below shows how popular models compare on cost. Artificial Analysis plots each one by its weighted average cost, in dollars, to run a single task on its Intelligence Index – lower is better.
This is big news because DeepSeek’s model is approaching the same performance level as Opus 4.8, but for a fraction of the cost. The same task costs $0.03 on DeepSeek’s V4 Flash and $3.15 on Claude Fable 5.
Competition between different labs is pushing the price point down.
And this all just happened this week. If we zoom out, the price of an AI token fell by more than 95% in three years.
But for most companies, the AI bill went up anyway.
OpenAI’s flagship was launched at $30 per million input tokens in March 2023. By late 2024, DeepSeek V3 shipped frontier-class output at $0.14, roughly one hundredth of that. Google now processes over 3.2 quadrillion tokens a month, up from 9.7 trillion two years ago, and the per-token price dropped the entire way.
So the model has continuously gotten less expensive. The line item for companies has not. Enterprise spend on large language models more than doubled in six months, from $3.5 billion in late 2024 to $8.4 billion by mid-2025 (Menlo research). Prices fell about 10x. Consumption rose far faster.
Prices are falling, but usage is increasing
Two numbers get quoted all the time – they point in opposite directions and both are true:
-
The price per token. Hold quality steady, and it’s collapsed.
-
The volume of tokens we’re burning. It’s exploded past the price drop.
Hold quality steady, and the price of a token has collapsed. Look at how much teams actually buy, and usage has blown right past it.
Google is the cleanest proof, because they showed both sides. At I/O 2026, Sundar Pichai said Google went from 9.7 trillion tokens a month to over 3.2 quadrillion in two years. That’s about a 330x jump. Over the same stretch, the cost of GPT-4-level quality fell from $30 per million tokens to under $0.50.
This isn’t new, it’s the oldest pattern in tech. When a better steam engine cut the coal needed per job, total coal use went up, not down. Cheaper steam just made it worth using for a lot more jobs.
Cheap tokens did the same thing to AI. The second a call got cheap enough, teams started making a hundred of them where they used to make one.
AI margins run lower than SaaS
This is where the token context impacts go-to-market.
In classic SaaS, a common pricing model was per seat, and the cost to serve one more user is often marginal. There’s still a cost of goods — hosting, storage, support — but it’s mostly fixed. It barely moves whether someone logs in once a month or a thousand times. That’s why software runs at 80-90% gross margin.
AI changes the shape of that cost. Every call burns tokens, and that’s a variable charge that grows with usage. Use the product twice as much, and that slice of COGS roughly doubles. There’s no “already paid for it.”
The numbers show it. Bessemer puts AI-native gross margins at 50 to 60%, against 80 to 90% for traditional SaaS. ICONIQ surveyed around 300 companies and pegged the average AI product margin at 52%, up from 41% two years ago. Two things follow.
First, the gap is closing on its own. As token prices fall, the same product earns more per dollar of revenue every quarter, without changing a thing.
Second, and this is the bigger one: a flat per-seat price can’t cover a variable per-call cost. Bolt an AI assistant onto an $80 seat, and inference can eat up quite a bit of margin.
That’s why seat pricing is breaking. The pricing model is now the lever that decides who keeps their margin.
What makes your AI bill grow
When a team’s AI spend triples, the token price is almost never the reason.
It’s usually one of four habits. Each one has a real number behind it.
-
Sending trivial work to the flagship model, at up to 50x the cost of a smaller one that could’ve done the job.
-
Agents that loop with no budget. One four-agent system looped for eleven days and burned $47,000 before anyone caught it. The dashboards looked healthy the whole time. Bloated context, re-sent on every step. The conversation grows, and the cost grows with the square of it.
-
No caching. Turn it on, and repeated input can run up to 90% cheaper.
The pattern underneath all four is the same.
The model isn’t necessarily expensive, the way it’s being used is.
How to protect your margins
The teams holding their margin do two things. They treat tokens like a real cost of goods. Then they pass that cost to the customer through pricing.
It starts with visibility. Tag every call by team and feature, so tokens show up as a COGS line you can actually see. You can’t cut what you can’t measure. From there, four moves on the cost side.
-
Route every call to the smallest model that can do the job. That alone cuts 40 to 85% of spend. Coinbase nearly halved its bill this way, and usage kept growing the whole time.
-
Cache aggressively. Repeated input runs up to 90% cheaper.
-
Cap hard. Set limits that actually stop an agent, not alerts that tell you after the money’s gone. A hard cap is what kills a $47,000 loop on day one.
-
Trim your context. Don’t re-send the whole conversation on every step. Cap the history, summarize old turns, and the cost stops growing with the square of the thread.
Then one move on the revenue side: reprice for the outcome. Put the variable cost on the customer, not your margin.
This is already happening. Salesforce Agentforce charges per action. Intercom’s Fin charges $0.99 per resolution, and nothing when it fails. GitHub Copilot moved to usage-based credits. And Paid.ai, built by Outreach founder Manny Medina with GTMfund among its backers, is building the billing rails for exactly this shift.
Tag @GTMnow so we can see your takeaways and help amplify them.
OpenAI’s CRO Denise Holland Dresser on why sales as we know it is over. The shift: from systems that record the business to systems that help run it. Customer signals are moving faster than pipeline reviews — the teams that win will be the ones that act on what they learn in real time, not the next reporting cycle.
Lantern research: AI shopping recommendations are far more stable than daily tracking suggests. Across 3,300+ AI responses, 80% of brand visibility remained unchanged over time — daily fluctuations reflect model variability, not real competitive shifts. Stop reacting to the dashboard noise and focus on the long-term factors that actually move the needle.
25 years of hiring GTM leaders — JD Peterson shares what actually works. The CMO at Perceptyx on why take-home assignments don’t predict good hires, and what he does instead: a live 30-60-90 day plan that forces sharper thinking throughout the whole process.
GTM: Inside Merge – How They Went Enterprise-First, Then Reversed to PLG
Listen through the links in the page above or by searching wherever you get your podcasts “The GTMnow Podcast.”
Proaction – raised $4.2M with GTMfund participating, to build the AI-native alternative to legacy fleet management. Fleets have relied on the same expensive, opaque vendors for decades. Proaction replaces call centers with AI agents and gives operators lower costs and real visibility into their fleet workflows.
Cathedral – raised $160M at a $1.4B valuation to build AI cyber weapons for the US military. Founded by four ex-DOGE staffers and led by a16z and Sequoia, it has no revenue and no announced contracts yet, so the price is a bet on Pentagon relationships and where defense-tech dollars flow next.
Glow – raised a $180M Series A at a $1.2B valuation to rebuild endpoint security for a world of AI. Born a unicorn out of stealth and co-led by Sequoia, Cyberstarts, Greenoaks, and Redpoint, the bet is that once AI usage on work devices tripled to 45% in a year, legacy tools can’t see shadow AI, plugins, or agents.
Fish Audio – raised a $52M seed to make voice the default interface for every AI model. A bedroom project trained on one GPU, now 8M users and $21M ARR, it clones a voice from 5 seconds of audio across 83 languages and is going straight at ElevenLabs’ economics.
Xsight Labs – raised $300M at a $2.8B valuation to build power-efficient chips for AI networking. Led by Fidelity with Intel Capital and Valor, its switch already powers SpaceX’s Starlink V3, and investors think AI data-center networking becomes a $150B market by 2028, taking share from Cisco and Broadcom.
Multiverse Computing – raised a $570M Series C at a $1.7B valuation to shrink AI models 80-95%. Co-led by Forgepoint, BNP Paribas, and Bullhound, its quantum-physics compression lets powerful models run on a laptop, phone, or factory floor with no cloud, a nearly 5x step-up from its round a year ago.
-
Senior Product Manager, Federal at Armada (Remote – Washington, DC)
-
VP, Revenue Marketing at CaptivateIQ (Remote – US / Toronto, Canada)
-
Account Executive, LatAm at Vanta (Remote – US)
-
Senior Lifecycle Growth Manager at Gorgias (Toronto, Canada)
-
Strategic Account Executive at Writer
See more top GTM jobs on the GTMfund Job Board.
Upcoming events you won’t want to miss:
-
Drive 2026: September 8-9, 2026 (Stowe, VT)
-
Lenny & Friends Summit: September 10, 2026 (San Francisco, CA)
-
Dreamforce 2026: September 15–17, 2026 (San Francisco, CA)
-
INBOUND: September 16–18, 2026 (Boston, MA)
-
Pavilion GTM2026: September 28–October 1, 2026 (NYC, NY)
-
Moment 2026: October 6, 2026 (NYC, NY)
-
CVC Week by Counterpart Ventures: September 29, 2026 (San Francisco, CA)
-
Customer Success Week: October 5-9, 2026 (NYC, NY)
-
TechCrunch DISRUPT: October 13–15, 2026 (San Francisco, CA)
-
GTMfund AGM/Retreat: October 15-17, 2026 (San Francisco & Napa, CA)
Some GTMnow Network love to close it out – we appreciate you.

















