Technology

VAST Data

A disaggregated all-flash data platform providing unified access via S3, NFS, and SMB protocols, optimized for AI and deep learning workloads with consistent low latency.

6 connections2 resources3 posts

Summary

What it is

A disaggregated all-flash data platform providing unified access via S3, NFS, and SMB protocols, optimized for AI and deep learning workloads with consistent low latency.

Where it fits

VAST Data targets the convergence of AI/ML workloads and object storage. Its all-flash architecture eliminates the cold scan latency that plagues spinning-disk object stores, while the S3 interface maintains ecosystem compatibility.

Misconceptions / Traps
  • Not just object storage. VAST is a unified data platform where S3 is one of multiple access protocols. Evaluating it solely as an S3 alternative misses its multi-protocol value.
  • All-flash means higher per-GB cost than HDD-based object stores. The value proposition is performance per dollar, not lowest cost per GB.
Key Connections
  • implements S3 API — S3-compatible interface
  • solves Cold Scan Latency — all-flash eliminates seek latency

Definition

What it is

A disaggregated, all-flash data platform providing S3-compatible object storage alongside file (NFS/SMB) and database access on a single unified architecture optimized for AI workloads.

Why it exists

AI/ML workloads need both high-throughput object storage and low-latency file access. VAST eliminates the need for separate storage systems by unifying all protocols on an all-flash platform with S3 compatibility.

Primary use cases

AI/ML training data platform, unified NFS+S3 storage, high-performance analytics, GPU-direct data access.

Recent developments

Latest signals
  • VAST Amplify — reclaiming stranded SSD capacity during the flash shortage (launched January 27, 2026). Program targeting the 2026 NAND supply crunch (30 TB SSD prices rose ~257% from mid-2025 to early 2026): existing NVMe drives get repurposed into E-Boxes/JBOFs, and DASE's global similarity-based dedup + erasure coding at only 12% overhead yields up to 6× effective capacity without new flash purchases. CEO Hallak has framed the shortage as "a wake-up call to the flash vendors" — VAST claims it would need flash supply growing 200% YoY to meet demand. Per VAST Data — VAST Amplify launch and Blocks & Files — VAST on the flash shortage.

  • The platform is now six engines on DASE — AgentEngine ships in 2026. Series F coverage detailed the "AI Operating System" build-out: DataStore, DataBase, DataSpace, DataEngine, InsightEngine, and AgentEngine — a production AI-agent deployment layer with MCP protocol integration — with PolicyEngine and TuningEngine slated by end of year. Storage vendor → agent-platform vendor is the trajectory to watch. Per Supercomputing News — VAST's $30B and the middle layer of AI.

  • Series F at $30B valuation — $1B round led by NVIDIA's AI-storage bet. Per VAST Data's press release on the Series F financing (April 22, 2026), VAST closed a $1B Series F at a $30B valuation. NVIDIA participation flagged across coverage as the marquee AI-infrastructure signal — VAST's positioning as the unified storage substrate for AI-training data loops is now backed by the GPU vendor with the most direct interest in storage stops being the bottleneck.

  • "Collapsing the Stack" strategy — own the AI data loop end-to-end. VAST's strategic frame for 2026 is consolidating the AI-data-pipeline tiers (object + file + database + lineage) onto a single all-flash platform — competing not just with object stores like FlashBlade and StorageGRID but with the broader "AI data lake + feature store + RAG infrastructure" surface that has historically been a multi-vendor patchwork. Per VAST — N=1: How VAST Built the De Facto Data Layer for AI.

  • The financials behind the $30B (June 2026 detail) + the CoreWeave anchor. Coverage of the Series F put hard numbers on the valuation: $4B+ cumulative software bookings, $500M+ committed ARR, ~90% gross margins, and 300%+ net revenue retention — plus a landmark $1.17B commercial agreement with CoreWeave that makes VAST the default storage supplier to GPU-heavy neoclouds. Per Calcalist — VAST valuation soars to $30B and Futuriom — VAST hits $30B.

  • From storage vendor to "AI Operating System" — native on BlueField-4. VAST's DASE (Disaggregated Shared Everything) architecture — stateless compute over a shared NVMe-oF flash pool — now bundles storage + a transactional database + a compute engine as one platform. Running a VAST CNode natively on NVIDIA BlueField-4 (STX reference architecture) enables zero-copy NVMe→GPU-VRAM paths via GPUDirect Storage + RDMA; VAST-reported benchmarks show a 90% inference-efficiency gain and a 20× Time-to-First-Token improvement when offloading KV caches to the tier. This is the same DPU-offload pattern the ROS2 ACM paper validates academically. Per VAST — N=1 data layer for AI.

  • The $30B software model has a name: "Gemini" — hardware and software are sold separately. Customers license VAST's software by storage capacity (TB/PB) and compute (CPU cores) on 4-5 year upfront terms, while hardware is purchased separately from partners at transparent at-cost pricing — the inverse of the classic storage-appliance bundle. Per TSG Invest — VAST Data Stock: $30B Valuation (tertiary source, no independent corroboration on the "Gemini" naming — flagged as such).

  • BlueField-4 KV-cache offload cuts power draw ~75% alongside the throughput gains. Independent research coverage of VAST's NVIDIA CMX (Inference Context Memory Storage Platform) integration puts a number on the efficiency side of the BlueField-4 DPU story already on the node: a 75% reduction in power consumption for KV-cache operations, on top of the previously reported 90% inference-efficiency gain and 20x TTFT improvement. Per NAND Research — VAST's Novel Approach to NVIDIA's CMX Platform.

  • AIC is building appliance hardware around VAST's DASE stack. AIC showcased its CERES Platform — NVMe storage built for VAST's DASE architecture — at NVIDIA GTC 2026, targeting AI data pipelines, RAG, and real-time analytics; a concrete sign of a hardware-partner ecosystem forming around VAST's software-defined model. Per AIC — AI Storage Platforms for Scalable Inference at NVIDIA GTC 2026 (tertiary, republished by multiple outlets incl. Yahoo Finance and Morningstar). Sources: VAST Data — VAST Amplify launch · Blocks & Files — VAST on the flash shortage (March 3, 2026) · Supercomputing News — VAST's $30B bet on the middle layer · VAST Series F at $30B valuation (press release) · VAST — N=1: The De Facto Data Layer for AI · Calcalist — VAST $30B / $1B raise · Futuriom — VAST hits $30B valuation

Connections6

Outbound4
implements1
Inbound2
competes_with2

Resources2

Featured in