When the Quota Became a Variable

Sometime this year, the effective size of our own $20/month Ollama Cloud plan changed. Earlier in 2026 we were getting roughly two million tokens through a session window; by September the same workloads on the same plan were hitting the wall around 700K ([user-measured] — our own first-party measurement, same plan and price throughout, dated 2026-09). No changelog entry. No email. No line item that moved. Just less.

That measurement is an anecdote, and this index has spent months building a standard that doesn't let us publish anecdotes as findings. So we did what the standard requires: treated our own grievance as an unverified claim, and pulled the documented record for the whole industry — changelogs, migration docs, archived pricing pages, dated news coverage — to find out whether the thing that happened to us is a pattern or a mood.

It's a pattern. It is also more interesting than the "rug pull" framing that dominates the community anger, because the verified record contains real expansions alongside the real squeezes. Both halves below, with every claim carrying its provenance — and a taxonomy we're going to hold ourselves to for the rest of this post.

Three registers of degradation

Any claim about AI subscription limits falls into one of three buckets, and mixing them is how bad takes get written:

  1. Confirmed silent degradation — effective capacity was reduced without announcement, and the reduction was later verified (vendor admission, refunds, or reproducible measurement).
  2. Announced change — the provider said it out loud, even if the framing was adversarial. Users may hate it; it is not silent.
  3. Unsubstantiated or unresolved community anger — a perceived reduction that can't currently be distinguished from model-mix effects, metering changes, or noise. Not fake, necessarily. Not verified either.

The Deep Research pull that seeded this post used roughly this taxonomy, and — in keeping with how we treat research inputs — we re-verified its classifications against primary sources before publishing. Several moved buckets. A few claims didn't survive at all; those are documented at the end, because a post about opaque metering that quietly dropped its own weak claims would be a self-own.

The documented ledger

Cursor, June–July 2025 — the pattern-setter (confirmed silent degradation). Pro's flat 500 "fast requests" quietly became a $20 pool of API-rate credits. Agentic, long-context workflows that fit comfortably in the flat quota started costing dollars per interaction; user reports of overage bills reached $350 in a single week ([user-measured]). The backlash forced a public apology from CEO Michael Truell and refunds for users charged between June 16 and July 4 ([vendor-published] — the apology and refunds are what upgrade the community's claims to confirmed).1 Every migration on this list since has followed the template Cursor established: keep the headline price, swap the unit underneath it.

Anthropic, August 2025 → September 2026 — a full cycle in fourteen months. The weekly cap on top of the rolling 5-hour session window arrived August 28, 2025 as an announced change targeting the heaviest users. In March 2026 came the one unambiguous confirmed-silent-degradation in Anthropic's record: peak-hour (5–11 AM PT) accelerated quota burn on paid plans, discovered when users measured Max session windows depleting in as little as ~19 minutes ([user-measured]) and confirmed by staff only after the backlash. Then a genuine reversal: on May 6, 2026, Anthropic doubled 5-hour rate limits and removed the peak-hour throttle entirely, on the back of a compute deal that gives it the full capacity of SpaceX's Colossus 1 datacenter in Memphis — 300+ MW, 220K+ GPUs ([vendor-published], widely corroborated).2 A "temporary" 50% weekly boost ran through the summer. And on August 29, 2026, Anthropic announced that boost ends September 14, replaced by a "permanent 25% raise" — which, computed against the pre-boost baseline, nets out to roughly 17% less weekly capacity than users have been running on through the boost period ([vendor-published announcement; net-reduction arithmetic verified in independent coverage]).3 Classification matters here: this is an announced change with adversarial framing, not a silent one — Anthropic said the quiet part in a changelog, and to its credit acknowledged the reduction when pressed. But it's the purest specimen we have of the quota-unit game: a cut, denominated as a raise.

GitHub Copilot — the canonical re-denomination, twice ([vendor-published], primary). Copilot went from "unlimited" flat-rate to 300 "premium requests"/month in 2025, with per-model multipliers that made the unit itself elastic — a base-model call cost 0×, a reasoning-model call 10×, and a single code review consumed 13 requests at once ([vendor-published multiplier tables, secondary compilation]).4 Then on April 27, 2026, GitHub announced the second re-denomination in eighteen months: premium requests deprecated entirely, replaced June 1 by GitHub AI Credits consumed at per-token API rates. Copilot Pro at $10/month now includes $10 in monthly credits.5 The flat-rate era didn't end with a price hike; it ended with the unit of purchase being redefined until it converged on metered API billing.

xAI, May 15, 2026 — deprecation as repricing ([vendor-published], primary migration doc). Grok 4.1 Fast — $0.20/M input, $0.50/M output — was retired along with seven other legacy endpoints. The mechanic is the important part: requests to the retired slug automatically redirect to grok-4.3 and bill at $1.25/$2.50 — roughly 6.25× on input — with no action required from the customer to start paying the new rate.6 We hit this one ourselves: our own routing config carried x-ai/grok-4.1-fast when the deprecation landed ([user-measured] — first-party). A deprecation with auto-redirect billing is a price increase that never has to be announced as one.

The proliferation tier. Vercel's v0 moved from fixed generation limits to variable token burn; its own community forum documents Premium credits exhausting on error-looped generations ([user-measured, community]).7 Perplexity Pro codified 300 searches/day ([vendor-published limit, aggregator-compiled]). Moonshot pulled its public Kimi Enterprise per-seat price card in August 2026, routing enterprise buyers to sales ([reported] — single aggregator source; treat accordingly).

Ollama Cloud — our own case, filed honestly (unresolved). Here is what we can verify about our own provider. The official pricing page, retrieved September 2026, meters the $20 Pro plan as "$60 of usage credits per month," "measured in tokens at each model's rates," resetting monthly, no rollover — and new Max sign-ups are currently paused ([vendor-published]).8 Third-party documentation of the same plans describes a different scheme: 5-hour session limits plus weekly limits reflecting "primarily GPU time," which depends on model size ([reported], third-party).9 Our ~2M→~700K session-yield measurement is consistent with either explanation — heavier default models burning an opaque GPU-time allocation faster, or the metering system itself changing under the plan mid-year. We could not pin the transition date from archived captures before publishing. So our own grievance gets filed in bucket three: unresolved, not confirmed. And that filing is itself the finding. A paying customer with logs, telemetry, and professional motivation cannot determine which of two metering systems governed their plan, or whether a 65% yield drop was a cut, a model-mix effect, or a re-denomination. When the unit is unobservable, real cuts and innocent drift are indistinguishable from the outside — which is precisely what makes the unit valuable to the seller. That's the AI Subscription Quota Opacity pain point this wave adds to the index.

The mechanics underneath

Four mechanisms recur across every case above, and they're worth naming because they're what the new node actually catalogs.

Unit re-denomination. "Active hours" (Anthropic), "premium requests" with floating multipliers (GitHub), "credits" (GitHub again, Ollama), GPU-time (Ollama's earlier vocabulary). Each abstraction moves the meter one step further from anything a user can independently count. Tokens are auditable — every gateway and observability tool counts them. Active-hours-of-milliseconds-while-the-model-processes are not.

Hidden reasoning burn. Inference-time reasoning models generate internal chain-of-thought tokens that bill against quotas users can't see or bound. We measured a GLM-family model burning ~16,000 reasoning tokens against a 628-token prompt ([user-measured] — our own measurement); the broader literature puts internal reasoning at commonly 10–30× the visible prompt-plus-answer volume ([independent-methodology and aggregator sources vary; the order of magnitude is corroborated, the precise multiplier is not standardized]). On a subscription metered in "active hours" or credits, that burn is invisible until the wall arrives.

Deprecation as forced upgrade. The Grok mechanism, in its purest form: retire the cheap slug, auto-redirect to the expensive one, bill at the new rate. The price of the service you were using went up 6×; the price of every listed product stayed the same.

Discretion in the terms. Here we must publicly drop a claim. The research input for this post quoted specific ToS language — Anthropic reserving the right to "accept, decline, condition, or limit any request… at our discretion" to prevent "unreasonable load" — with citations that turned out to resolve to unrelated third-party websites' boilerplate. We fetched Anthropic's actual Commercial Terms and did not find the quoted clause; the nearest real provision is a security-grounds suspension right. So the quote is out. What survives verification is weaker but still meaningful: the March 2026 peak-hour throttle operated for weeks inside whatever contractual room exists, and no plan document enumerated it. The discretion is real; the dramatic quote wasn't.

The counter-evidence, printed at full size

A thesis that can't survive its counter-evidence is a mood. Here's the strongest material against "the great AI rug pull," and it's substantial:

  • OpenAI uncapped text. On August 6, 2026, message caps on standard text conversations were removed across every ChatGPT tier including Free ([aggregator-consistent reporting]; the same sources note that current-generation reasoning modes now trade time rather than quota, replacing the older per-week reasoning caps).10 That is a genuine, large expansion of baseline utility — with the important asterisk that Deep Research, agent mode, image, and voice remain firmly metered. Commodity text became free; premium compute became the product.
  • Anthropic's May 6 doubling was real, announced, and reversed a genuine degradation ([vendor-published]).2 Fourteen months of Anthropic history contains one silent throttle, one doubling, one temporary boost, and one adversarially-framed net cut. That's a provider oscillating around brutal unit economics, not a monotonic squeeze.
  • The floor keeps falling. DeepSeek made its 75% V4-Pro cut permanent in May 2026 — $0.435/M input, $0.87/M output as the standing list price, $0.14/$0.28 for V4-Flash ([vendor-published]; these are the figures already verified on our DeepSeek V4 node, and they differ from some aggregator-repeated "off-peak" figures in circulation).11 Zhipu's GLM-Flash class remains free. Whatever is happening to Western flat-rate subscriptions, per-token frontier-adjacent inference has never been cheaper.

So the honest thesis is not "prices went up." It's bifurcation plus unit games: commodity inference is being pushed toward free while premium/agentic compute is being metered in units the buyer cannot audit — and the transition between those two regimes is where the documented degradations live.

Why it's happening: the economics nobody has to whisper

The COGS story requires provenance labels more than any other section, because its most-quoted numbers are the least verified. Analyst estimates place provider losses on subscription power users at 10–40× revenue, with the viral figure — a single $20 account consuming ~$14,000/month in API-equivalent compute — attributed to SemiAnalysis but circulating mostly through secondary channels ([aggregator-repeated]; treat as directional).12 The counter-number, from leaked Q3 2025 OpenAI financials, claims $1.75 of revenue per $1 of inference cost — but analysts promptly noted it appears to exclude CapEx amortization ([aggregator-repeated], both directions).13 What's solid is the anchor underneath: on the order of $400B/year is going into AI infrastructure that must eventually be paid for by inference revenue ([independent-analyst estimate]).14 Flat-rate plans in front of that CapEx wall are loss-leaders by construction, and the lifecycle — loss-leader, dependency, squeeze — is the oldest curve in SaaS, run this time on the most expensive COGS in software history. The squeeze isn't a conspiracy. It's arithmetic arriving on schedule. The Vendor Lock-In node has always said dependency is the precondition; 2026 is the year the invoices agreed.

The exit this index has been mapping all along

Here is the part we get to say with the receipts of priority: the High Cloud Inference Cost pain point has been on this index since its earliest waves — mapping per-token API cost as the binding constraint on LLM-over-S3 workloads long before this September's subscription news. The quota-opacity wave is that pain point's subscription-tier sibling, and the solution edges were already drawn:

Measure. Gateway observability — LiteLLM, Helicone-class tooling — translates every opaque unit back into token counts and dollars per virtual key. If your inference spend matters, the meter should be yours, not the vendor's.15

Route. Multi-provider gateways exist precisely so that a deprecation like Grok 4.1 Fast's is a config change, not a 6× bill. OpenRouter aggregates 400+ models behind one API; Portkey-class managed gateways add fallbacks and budget-capped virtual keys ([vendor-published category descriptions]).16 Routing turns provider pricing moves from binding constraints into inputs.

Own the floor. The Local Inference Stack is the terminal answer, and the one this whole index is oriented around: open-weight models (DeepSeek V4 ships 1.6T-parameter open weights under MIT; Qwen, GLM, Llama fill the ladder below) on owned hardware, served by vLLM for concurrency or Ollama for the single-dev tier, against data that never leaves your object store. We'll be straight about the TCO literature: the published 2026 crossover analyses are aggregator-grade and internally inconsistent (we found the same RTX 5090 workstation quoted at two different prices and the wrong VRAM inside one analysis), so we won't reprint their request-per-day thresholds as fact. The structural claim doesn't need them: amortized owned-hardware cost is flat and auditable; subscription effective-cost is variable, rising for agentic workloads, and metered in units you can't see. A $2–3.5K workstation amortized over three years is a known number every month. No one re-denominates your GPU.

The bifurcation the counter-evidence revealed is, in the end, the same two-tier architecture our DeepSeek coverage has been describing since June: an open, cheap, self-hostable floor doing the token-heavy work against the data plane, and a metered closed frontier reserved for the reasoning peaks. The subscription squeeze doesn't argue against that architecture. It prices it in.

Our read

Dated call, 2026-09-03: by end of 2027, (a) effective-quota observability becomes a named product category — subscription-quota reconstruction shipping as a first-class feature in gateway/observability tools, because no US frontier provider exposes an auditable quota API; (b) at least one major US provider publishes a machine-readable usage-accounting surface for subscription plans under competitive pressure; and (c) the two-tier exit — open-weight floor for volume, metered frontier for peaks — becomes the documented default in enterprise AI architecture guidance, driven by quota risk at least as much as by capability or price. We revise if providers converge back to transparent flat quotas, if the open-weight floor reprices upward (DeepSeek-class permanent cuts reversed), or if churn evidence shows users simply absorb the opacity and the countermeasure categories stall. Logged on the calls ledger.

What changed on the index

One new pain-point node: AI Subscription Quota Opacity — the case ledger above, with classifications and counter-evidence, now lives in the graph with solves edges from Local Inference Stack, LiteLLM, and Helicone AI Gateway. Enrichments landed on High Cloud Inference Cost (the subscription-tier squeeze block), Ollama (the Ollama Cloud metering-vocabulary record), and Local Inference Stack (the squeeze as adoption driver). One new dated call on /calls. And for the record of what verification removed from the research input before publication: a future-dated event reframed to its actual announcement date and classification; two fabricated-citation ToS quotes dropped after checking the real terms; a $15 credit figure corrected to $10 against the primary source; current-state OpenAI reasoning caps re-dated as historical; DeepSeek pricing reconciled to our already-verified node figures; and a TCO table set aside for internal contradictions. The index stands at 438 nodes.

Works cited

Footnotes

  1. Solvimon — what Snowflake, Twilio, and Cursor's pricing rollouts teach about token economics — the Cursor June 2025 timeline, backlash, apology, and refunds.

  2. Appwrite — Anthropic just doubled Claude Code's 5-hour rate limits — the May 6, 2026 doubling, peak-hour throttle removal, and SpaceX Colossus 1 compute deal. 2

  3. BleepingComputer — Anthropic is cutting Claude Code's current weekly limits by 17% — August 29, 2026 coverage with the boost-vs-baseline arithmetic; effective September 14.

  4. Olumia — GitHub Copilot premium requests: allowances, costs and limits — the multiplier tables and the 13-requests-per-code-review mechanic from the premium-request era.

  5. GitHub — Copilot is moving to usage-based billing — primary announcement, April 27, 2026: premium requests deprecated for AI Credits effective June 1; Pro $10/month includes $10 in credits.

  6. xAI docs — Grok model retirement on May 15, 2026 — primary migration doc: eight legacy endpoints retired, automatic redirect to grok-4.3 at $1.25/$2.50 per M tokens.

  7. Vercel community — v0 Premium: credits gone fast — user-measured credit exhaustion after the token-based pricing shift.

  8. Ollama Cloud — plans and pricing — vendor-published current state, retrieved September 2026: Pro $20/month with $60 usage credits, token-metered at per-model rates, monthly reset, Max sign-ups paused.

  9. Ollama TPS — Ollama Cloud limits explained — third-party documentation of the session-plus-weekly scheme metered "primarily in GPU time."

  10. AI Toolbox — ChatGPT limits 2026 — aggregator compilation of the August 6, 2026 text uncap and the shift from reasoning quotas to effort-time tradeoffs.

  11. DeepSeek API docs — models and pricing and InfoWorld — DeepSeek's steep V4-Pro price cut escalates the AI pricing war — the permanent 75% cut and standing rates.

  12. r/BetterOffline — SemiAnalysis on serving Opus-class models at $20 — aggregator-repeated attribution of the ~$14K/month power-user figure; directional, not primary.

  13. r/BetterOffline — OpenAI losses reporting — the leaked $1.75-per-$1 inference figure and the CapEx-exclusion critique; aggregator-repeated in both directions.

  14. LessWrong — Power Overwhelming: dissecting the $1.5T AI revenue shortfall — independent analysis of the CapEx-vs-revenue gap underneath subscription economics.

  15. TrueFoundry — Claude Code rate limits and usage quotas explained — secondary compilation of the active-hours mechanics and the observability counter-move.

  16. Respan — OpenRouter alternatives for LLM apps in 2026 — the multi-provider gateway landscape as deprecation/limit hedge.