This index has tracked the Catalog-Centric Control Plane since June 2026, when Polaris, Unity Catalog and Gravitino converged on the same idea: the catalog stops being a lookup table and becomes the place where access decisions get made. September 2026 completed the thought and exposed its cost. The catalog is now the component that decides who gets credentials to your object storage — and this month produced one release that made catalogs agent-callable by default, one that broke every long-lived integration to shorten credential lifetime, and one CVE showing what happens when the layer holding a vended credential leaks it.
The pain point underneath all three is Credential Vending, and it is worth stating plainly before the releases: vending short-lived credentials narrows the blast radius only if every consumer between the catalog and the user treats them as secrets. A query engine is a consumer that hands users a programmable interface by definition.
The catalog became a tool an agent can call
OpenMetadata 2.0.2 landed September 16 and does not describe itself as a data catalog any more. The project now positions as "The Open Context Layer for Data and AI," and the architecture follows the tagline rather than decorating it.1 The metadata store is restructured as a semantic context graph with RDF support — a traversable network rather than a set of rows — and the Model Context Protocol is enabled by default, which turns the catalog into a suite of callable tools. An agent can run a semantic search, walk lineage, and mutate metadata at query time.2 A new AI Studio and Context Center ingest unstructured operational documents — runbooks, wiki pages — straight into the metadata graph, attaching hybrid vector and keyword search to individual assets.3
Read that as an access-control statement rather than a feature list. The catalog was already the place that knew where every table lived and who could touch it. Making it agent-callable by default means a class of non-human caller can now traverse and modify that knowledge, and the default is on.
The migration cost is real and under-advertised. OpenMetadata's own breaking-changes documentation instructs operators to take a full database backup before migrating, because there is no automated downgrade: the 2.0.0 script creates 15+ new tables, renames thread_entity, and rewrites application, service-connection and tag rows in place.4 Four more changes bite in production. The Airflow-based internal orchestrator is deprecated in favour of the native Kubernetes Orchestrator introduced in 1.12. The thread-backed collaboration model is rewritten onto a new Tasks entity, and the /v1/suggestions endpoint is removed outright. Embedding and natural-language-query settings move out of elasticsearch.naturalLanguageSearch into a dedicated llmConfiguration block. And custom ingestion images must be rebuilt against Python 3.12 — CP310 binary wheels fail to import on the new base image.5 Anyone running pinned ingestion containers should budget this as a rebuild, not an upgrade.
One thing 2.0 conspicuously did not do: compete. The release shipped a Unity Catalog model-version resolution fix and a native Apache Polaris credential-vending record, but no Iceberg REST catalog surface of its own.2 The posture is bridging — sit semantically above Unity Catalog, Polaris and Gravitino rather than replace them. For the consolidation question this index has an open call on, that is evidence against near-term convergence on any single control plane. The metadata layer is betting the catalog war stays unresolved long enough for a semantic layer to be worth buying.
The token that used to last forever
Unity Catalog OSS 0.6.0 contains the highest-blast-radius change of the month, and it is one line in the release notes.
Before 0.6.0, access tokens issued by the /auth/tokens exchange had no expiration. They were valid indefinitely. Under 0.6.0 they expire on a server-dictated default of 24 hours, governed by server.access-token-timeout.6 Every client that stashed a token once and assumed it would keep working — embedded connectors, CI jobs, the script someone wrote in 2025 — now needs re-exchange or refresh logic. Nothing fails at upgrade time. Things fail a day later, which is the worst possible shape for an operational surprise.
The same release re-keyed credential caches on the full vending context — catalog endpoint, storage scheme, and caller authentication — rather than on table path alone.6 That sounds like bookkeeping and is not. A cache keyed only on path is the shape that lets one caller's vended credential satisfy another caller's request. Keying on the whole context closes it. Taken together, the two changes are a single argument: shorten how long a credential is valid, and tighten who a cached one can be handed to.
0.6.0 also implements Metric views on Apache Spark 4.2.x and promotes SQL VIEW objects to native catalog citizens, storing raw SQL text, column lists and read lineage for cross-engine resolution.6 Combined with OpenMetadata's semantic-context push, the pattern is consistent across both releases: the semantic layer is migrating down into the catalog.
What it looks like when the boundary leaks
Trino is where the abstraction failed, and the CVE reads like a specification of the risk the other two releases are managing.
From version 439 up to (but excluding) 480, the Iceberg connector's REST catalog static credentials (access keys) and vended temporary credentials were accessible to users holding only SQL-level write privilege.7 That is a privilege inversion. The query engine was handing out the keys to the bucket underneath it, so an account scoped to write a single table could extract credentials reaching the whole storage layer — and then bypass Trino, its row filters, its column masks and its audit log entirely. Patched in 480. The scoring is split: NVD rates it 7.7 HIGH while the CNA assigned 6.5 MEDIUM, both CVSS 3.1. Read the higher score as the one assuming a realistic multi-tenant deployment.
Note what did not go wrong. The vending layer worked as designed — it issued a scoped, temporary credential exactly as Credential Vending describes. The failure was one layer up, in the consumer that held the credential and exposed it to the SQL surface. Short-lived credentials limit how long a leak is useful; they do not stop the leak.
The other September: formats got faster, quietly
Away from the credential boundary, the storage layer shipped real mechanism.
Lance Format 2.1 was marked stable in Lance 0.38.0, and LanceDB published how the 1–2 IOPS random-access guarantee actually works.8 The format splits the problem by value size. Mini-block handles small values, packing them into indivisible 4–8 KiB blocks — one to two disk sectors — using opaque encodings such as delta: maximum compression, slight read amplification, and crucially no per-value reverse shredding in memory. Full-zip handles large values like dense vectors and raw images, where mini-blocking would fracture a single value across blocks and destroy read performance; it compresses a whole 8 MiB page with strictly transparent compression, accepting in-memory zipping to recombine buffers.9 Both lean on a lightweight repetition index mapping row offsets straight to block file offsets, which is what lets Lance skip the recursive scan Parquet's definition levels require. Format 2.2 extends it: run-length encoding promoted from mini-blocks to full block-level, General Block Compression applying LZ4 automatically above 32 KB, and a native Map type so nested map values keep the IOPS guarantee and fields can be added or removed without column rewrites.10
The post-Parquet story also got its first named production evidence. Our 2026-07-21 call asked for independent case studies converting vendor benchmarks into deployment reality. LanceDB now publishes them: Midjourney's CFO on Lance carrying the vector infrastructure behind their generative pipelines, and Runway's Head of Engineering crediting column-append-without-rewrite plus fast random access with reshaping their model-training pipelines.11 These are vendor-hosted testimonials from named executives — reported, not independently audited — which is a real step up from a benchmark chart and still short of a third-party writeup. The clause is confirmed for Lance and not for Vortex or Nimble, which still have no named third-party deployments. Scoring one format's evidence as the whole field's would be the easy mistake.
ClickHouse 26.8 LTS is a backward-incompatible release for object-storage disks: operations now use the metadata storage's native transactions rather than the previous simulated ones.12 That is a correctness-path change to the layer deciding whether a partially-written mutation is visible — exercise it on a copy first. The same release adds run_query_in_background, which accepts a query, returns an empty result immediately, and runs it to completion regardless of whether the client connection drops — aimed squarely at large INSERT ... SELECT pipelines against lake storage, where the job outlives any reasonable client timeout.13
OpenDAL 0.59 is an API overhaul rather than a version bump: Metadata is constructible only through MetadataBuilder, every direct constructor and with_* method stripped; DeleteInput removed outright in favour of a (String, DeleteOptions) tuple; OpCopier merged into OpCopy.14 SeaweedFS 4.47 is an operations release — a visual IAM policy editor, an integrated Rust maintenance worker shipping alongside the Go binaries, S3 API retry logic, and volume files opened with O_NOATIME to cut metadata write amplification.15
And a negative result worth recording, because adjacent progress is not the same as the thing: vLLM 0.28–0.29 shipped substantial KV-cache and cold-start work — a persistent per-GPU weight-cache daemon ("Fast Start") mapping post-quantized TP-sharded weights over CUDA IPC via --load-format ipc_cache, Mamba prefix caching at a measured 9–25% TTFT improvement, and Model Runner V2 as the default with CUDA-graph profiling for KV-cache auto-sizing.16 None of it offloads to object storage. The memory-tier work is real and stops at GPU and host memory. The S3 tier remains LMCache and Mooncake territory, and our standing call on a billable KV-cache-on-object-storage product stays unmoved.
The ontology delta
No new nodes this wave — every concept in the source research already had one, which is its own signal about coverage. Ten existing nodes gained substance: OpenMetadata, Unity Catalog, Credential Vending, Trino, ClickHouse, OpenDAL, Lance Format, LanceDB, SeaweedFS and vLLM. The index holds at 438 nodes.
Two corrections are filed on the calls ledger. Our Trino node claimed release 481 when upstream had shipped 483 two months earlier — it survived the daily staleness sweep because Trino's bare-integer versions cannot be parsed into a comparison, so the node was reported as unknown rather than stale. A verification instrument's coverage is a separate fact from its findings, and only the findings make the headline. And the four standing calls re-checked this month produced one CONFIRMED and three NO NEW DATA, recorded with equal weight, because a ledger that only logs movement is a ledger that flatters itself.
Footnotes
-
The Open Context Layer for Data and AI — OpenMetadata on GitHub ↩
-
CVE-2026-34214 — NVD. Published March 31, 2026; patched in Trino 480. ↩
-
Columnar File Readers in Depth: Structural Encoding — LanceDB ↩