OpenAI priced AI by the person, DeepSeek by the hour of the day, Meta at zero on hardware you already own, and two vertical tools by what the employee they replace earns, which means per token and steadily falling is no longer a safe thing to budget on.
Every AI budget written in the last two years rests on one assumption nobody bothers to write down: the meter runs per token, and the number on it falls over time. This week five separate vendors retired that assumption, and not one of them replaced it with the same unit.
Start with the two that changed the shape of the bill rather than the size of it. OpenAI introduced a Premium seat for ChatGPT Business at $125 a month, or $100 on annual billing, sitting on top of a $25 Standard seat that is not going anywhere. Five times the price, five times the usage, no five-hour window, and workspaces can mix the two. That is a company saying out loud that not every user is worth the same amount, and that the ones worth the most were already expensing personal accounts to get around the cap. Then DeepSeek split its rate card by hour of the day, with off-peak set at half of peak and everything going up. The number worth staring at is cached input, the cheapest line on the old card and the single largest one in any agent that runs more than one turn, up more than twelvefold at peak.
Google looked like the counterexample and mostly is. Gemini 3.7 Flash landed Wednesday at $0.75 per million input tokens and $3.75 output, roughly half what its predecessor originally cost, with real gains on the benchmark that measures business workflow completion rather than trivia. Read the fine print and the discount has a date on it. It expires December 31 and the rate doubles on January 1. So even the good news this week arrived structured as a promotion rather than a floor, which is a different thing to build a two-year plan on.
The second group moved the meter onto hardware you already bought. Meta open-sourced Muse Glimmer, a 30-billion-parameter agent model under Apache 2.0 that fits in 24 to 32 GB quantized and runs on a MacBook or one consumer GPU, listed at $0.00 per million tokens for the obvious reason that you are the one running it. Cloudflare came at the same problem from the compute side with Kitesurf, a browser engine with no Chromium underneath it, running agent work at about a third of the CPU and a seventh of the memory. Both attack the identical line item, which is the reason your ops team still has a person doing eleven portals every Monday. That project was never blocked on intelligence. It was blocked on the fact that watching something continuously has always cost the same as thinking about it once.
The third group did something more interesting than repricing tokens. It priced against payroll. GIGR's Playad Autopilot sells the whole paid media loop on a page where the tiers are named Intern, Part Time, Full Time, and CMO, and the Full Time plan is $499 against the $3,000 to $8,000 a mid-market advertiser pays an agency to operate the same spend. An insurance quoting agent shipped at 1,050 quotes inside a $499 plan, which divides out to 48 cents a quote, or about one minute of a licensed producer's median wage. A free MIT-licensed pack of fourteen agent skills covers the scope a fractional CMO bills $200 to $350 an hour for. Nobody in this group is comparing themselves to software anymore. They are comparing themselves to a person, in that person's own units, on a public pricing page.
And the fourth group moved cost off headcount entirely. PPT Master hands back a real editable .pptx on your existing brand template against Microsoft 365 Copilot's $21 per user per month. Macro puts email, chat, docs, tasks, and a CRM in one AGPLv3 codebase against roughly $51 a seat for Slack, Linear, and Notion. A Claude Code skill for editorial diagrams went at the Visio and Lucidchart and Miro seat licenses your team renews without reading. Remix lets a product manager ship a live variant of the real app for $29 to $99 a seat plus agent credits. Same argument in four costumes: stop paying by how many people might touch this, start paying by how much of it actually gets done.
Two stories this week were not about price and still set the ceiling on all of it. Anthropic began watermarking Claude's text worldwide, at the model level, with no surface you can route around, which turns every "original, human-written content" clause in your vendor contracts from theater into something one side can substantiate. And the week's lesson on the three permissions you should never stack on one assistant is the standing reminder that a cheaper agent you cannot safely give private data, untrusted input, and an outbound channel is not actually cheaper. It is just unbuilt at a lower price.
Now the part the week got wrong. Open source eating the per-seat subscription was the loudest theme on the board and it carried the least honest arithmetic, and you do not have to take my word for it, because three of those projects said so about themselves. PPT Master's own honest section tells you to compare it to token spend rather than to zero, since it wants a long-context model and image keys and an agent environment to drive it. The go-to-market skill pack notes there is nobody to call. Macro's self-hosting story gets appropriately uncomfortable when you ask who runs the Docker stack. Every one of these moves the bill from a published invoice to an unpublished token meter plus somebody's afternoon, and only the invoice was ever on a page. The under-discussed version of the same problem sits on the commercial side: Remix, Playad, and the insurance tool all publish a tidy seat or plan price and then meter the actual work in credits with no published conversion rate. Four of this week's tools bill in a currency they do not define. That is the real repricing, and it did not make a single headline.
If you only have time for one thing this weekend, make it the DeepSeek repricing. It is the only story in the window that invalidates a spreadsheet you have already written and shown to someone. Cached input going up more than twelvefold hits multi-turn agents hardest, which is to say it hits the workloads that looked cheapest. Convert the peak windows and they land on the Chinese working day, so a team on ordinary American business hours never touches peak and is looking at the off-peak column instead, output up about 128 percent. The exception is the one to check today: if any part of your pipeline runs on an overnight batch, it now sits squarely inside the expensive window, and moving it is a cron change rather than a project.
Inference spent two years behaving like a falling number you could look up. This week it started behaving like a utility bill, a seat license, and a wage, depending on who was selling it. The unit is the decision now, and somebody else is making it while you are still comparing dollars per million.