Storage for inference
Object storage stopped being where you put things and became where models think. That is the argument this index has been making since March 2026 — not in one essay, but across seven posts as the evidence arrived. Read in order, they are a single line of reasoning.
Register: our read. Each step links to the post that made it, with that post's own summary. The underlying facts live on thenode pages; the editorial calls, with their revision conditions, live in the calls ledger.
The premise — inference starts moving to where the data already is.
The Shift to Local Intelligence: How Local S3 and Edge Inference Are Replacing Cloud Latency
The economics and architecture behind the migration from cloud inference to local-first AI — covering the Local S3 semantic storage pattern, CXL memory expansion, specialized inference silicon, and why 80% of AI inference is moving to the edge.
The consequence — agents that hold state need storage that holds state.
When AI Memory Became an Architecture: KV-Cache Persistence, MCP, and the Night S3 Got Its Memory Tier
Stateful agents needed stateful storage. Tonight 33 new nodes joined the index: Mem0, Zep, LMCache, SGLang, Mooncake, MCP, cuObject, BlueField-4, Animesis CMA. The architectural thesis: AI memory infrastructure crystallized around object storage in 2026, and the persistence layer the industry settled on is S3.
The response — the industry answers a constraint this index had tracked for months.
Answering the Memory Wall: 40 New Nodes Across the May 2026 AI Memory Inflection
The Memory Wall has been a named pain point on this index for months. The May 2026 inflection is the industry's answer in four cohesive movements — memory governance, KV-cache mechanics, the MCP orchestration layer, and durable agent runtimes. This post is the field report on the 40 new nodes added to the index over the past 48 hours: what each cluster covers, what they share, and why the answer surface had to land all at once.
The cost — persistence brings a security class that statelessness did not have.
When Memory Became the Attack Surface: The May 2026 AI Agent Security Inflection
Prompt injection was stateless. Memory poisoning is persistence. The OWASP Top 10 for Agentic Applications classified ASI06 in early 2026, BEAM 10-million-token benchmarks exposed the limits of context expansion, and the defense substrate moved into the memory layer. This index now maps that inflection across six new nodes.
The hinge — the cost of intelligence falls below the cost of remembering.
When Inference Became Cheaper Than Storage: The May 2026 Cost Inversion
In May 2026 DeepSeek made its 75% price cut permanent. V4-Flash output is now 107x cheaper than GPT-5.5. For the first time the cost of intelligence dropped below the cost of moving data — and every storage decision made for the 2024 cost curve is now suspect. This is the continuation of the I/O-stack rewrite.
The market — vendors, benchmarks and CVEs arrive, as they do once something is worth money.
Memory Became a Market: The AI Data Stack is Being Rebuilt on Three Fronts
In one week of July 2026: a benchmark war between agent-memory vendors with irreconcilable numbers, a second critical CVE in the most popular memory layer, a columnar format field fracturing into workload-specific successors to Parquet, and 3-bit KV-cache quantization getting its first independent reproduction. Underneath all three fronts, a runaway enterprise SSD price surge is forcing the whole stack onto object storage. The rebuild isn't coming — it's live.
The arithmetic — a break-even number for persisting inference state to object storage.
The Storage Tier Is the New Context Window
A vendor published the break-even arithmetic for persisting LLM inference state to object storage: one cache reuse every 94 hours justifies the spend. Between that number, NVMe object tiers priced like commodities, and streaming engines that write payloads straight to S3, the quiet reclassification of object storage — from cold archive to active memory tier — now has receipts.