WEKA
WEKA is an AI-native, software-defined parallel storage platform (WekaFS) that presents a single namespace across NVMe flash with S3, POSIX, NFS, and SMB access. In 2026 its center of gravity shifted from "fast filesystem for AI training" to **inference-memory infrastructure**: the **Augmented Memory Grid** pools NVMe across nodes via a custom user-space RTOS and RDMA, and offloads the LLM **KV-cache** from GPU HBM to that persistent tier with sub-millisecond retrieval — a measured **7.5 million read IOPS**. It is one of the principal vendors defining the [Inference Context Memory Storage](/node/inference-context-memory-storage-icms) ("Tier 3.5") layer alongside VAST, Dell, and HPE.
Definition
WEKA is an AI-native, software-defined parallel storage platform (WekaFS) that presents a single namespace across NVMe flash with S3, POSIX, NFS, and SMB access. In 2026 its center of gravity shifted from "fast filesystem for AI training" to **inference-memory infrastructure**: the **Augmented Memory Grid** pools NVMe across nodes via a custom user-space RTOS and RDMA, and offloads the LLM **KV-cache** from GPU HBM to that persistent tier with sub-millisecond retrieval — a measured **7.5 million read IOPS**. It is one of the principal vendors defining the [Inference Context Memory Storage](/node/inference-context-memory-storage-icms) ("Tier 3.5") layer alongside VAST, Dell, and HPE.
WEKA is a load-bearing example of the thesis this index tracks — object/flash storage becoming an **active extension of GPU memory** rather than a cold archive. By letting serving engines page KV-cache to a shared RDMA-attached NVMe tier instead of pinning it in ~$100/GB HBM, WEKA lets a cluster hold context for multi-turn agentic sessions across days while keeping GPUs saturated. Its S3 interface makes that tier interoperable with the rest of the object-storage ecosystem rather than a proprietary island.
KV-cache offload / persistence for high-concurrency LLM inference; AI-training data delivery that keeps GPUs fed; RAG and agent-memory backends needing sub-ms retrieval at scale; mixed POSIX+S3 pipelines on one namespace.
Recent developments
- NeuralMesh AI Data Platform GA — built on the NVIDIA AI Data Platform reference design (GTC 2026, March 16). WEKA's appliance-style AIDP system packages the stack for AI-factory deployments, cutting deployment "from months to minutes" and claiming 6.5× more tokens per GPU for inference via the Augmented Memory Grid; launch partners include Red Hat and Neysa. This is the NVIDIA-reference-design procurement play arriving in the file/KV tier. Per HPCwire — WEKA Releases NeuralMesh AIDP and WEKA — Turnkey NVIDIA AI Data Platform Solution.
- OCI validation puts hard numbers on the Augmented Memory Grid. On Oracle Cloud bare-metal H100, WEKA + OCI report 10× concurrent users vs DRAM-only (5,000+), ~2M tokens/sec, 7× tokens served in stress tests, ≥320 GB/s read / 150 GB/s write, and 4–20× TTFT improvement. Delivery stack for NeuralMesh AIDP: Supermicro hardware, Spectro Cloud PaletteAI one-click Kubernetes, an open-source vector-DB layer, and prebuilt vertical pipelines. One independent caveat (Lockwood): WEKA has no native BlueField-4 client-side integration, so in a strict CMX topology it participates as an external tier rather than on the DPU itself. Per WEKA + OCI 10× validation (PRNewswire) and NAND Research — NeuralMesh AIDP/STX.
- Scality partnership deepened (July 8, 2026) — NeuralMesh compute tier + RING object tier. Expanded from the February 2026 validated joint solution (WEKA NeuralMesh paired with a Scality RING object tier — up to 10× faster performance, ~20% lower infrastructure cost) into a distribution + joint-support agreement announced at RAISE Summit. The hot-tier/capacity-tier pairing is becoming a named product, not a reference PDF. Per StorageNewsletter — Scality and WEKA deepen partnership.
- Augmented Memory Grid — KV-cache as a persistent, RDMA-attached tier. WEKA's 2026 "Context Era" positioning pools NVMe across nodes and serves KV-cache offload with sub-millisecond latency and ~7.5M read IOPS, letting AI serving engines maintain multi-day agentic context without consuming constrained HBM — the same ICMS / KV-cache-disaggregation pattern productized. Per WEKA — The Context Era Has Begun and How WEKA is Solving AI's Trillion-Dollar Memory Problem.
- Cloud-provider integrations for inference. OCI × WEKA coverage frames the grid as reshaping AI-inference infrastructure at hyperscaler scale, extending the tier beyond on-prem. Per OCI and WEKA reshape AI Inference Infrastructure (AI CERTs).
- Training-side framing: GPU performance depends on storage. WEKA's own engineering material makes the case that the training bottleneck is the data/context path, not raw FLOPs — the demand-side argument behind the whole GPUDirect/RDMA-to-object-storage build-out. Per WEKA — AI Training: GPU Performance Depends on Storage.
Connections 9
Outbound 9
scoped_to2implements1augments1competes_with1