Ai Engineering 4 min read

Turbopuffer Retires Its Vector-Primary Architecture: 'RIP, Vector Database'

Turbopuffer announced v3 of its search engine on September 30, demoting the ANN vector index from primary key to just another secondary index, after serving 100B+ vector indexes at 200ms p99; the post argues vector-primary storage has been pushed as far as it can go.

Turbopuffer, the search infrastructure behind a meaningful share of AI retrieval workloads, published a post on September 30 titled “RIP, vector database”, and the title undersells the substance: the company is rearchitecting its entire engine in v3 to demote the ANN vector index from the primary key everything hangs off to just another secondary index. The current system it is replacing is no toy: single indexes of 100 billion-plus vectors, 200ms p99 reads at 1,000+ queries per second, over object storage. When the vendor that pushed that architecture furthest declares it a ceiling, the engineering reasoning is worth reading in full.

The Three Problems With Making Vectors the Primary Key

The v3 rework is driven by failure modes that only appear at scale. Storage amplification: multi-vector documents (late-interaction and ColBERT-style embeddings) duplicate the full document content once per vector, so a document with fifty embeddings is stored fifty times. Write amplification: SPFresh’s rebalancing moves vectors between clusters, and under vector-primary layout each move drags the full document and every referencing inverted index along, in the post’s words, “updating just one vector can move hundreds of attributes and their indexes.” And limited vectorization: ANN clusters of roughly 100-200 documents constrain block sizes for every query plan, even though the rest of the database world long ago settled on larger batches (DuckDB uses 2,048-row batches, ClickHouse around 65k, Lucene 256-doc posting blocks). The proof that decoupling pays came first in full-text search: reworking FTS v2 postings to fixed 256-blocks made the index 10 times smaller and queries up to 20 times faster, and v3 applies the same decoupling to vectors.

The Rename Is Half Marketing, Half Correct

“RIP, vector database” is a claim about architecture, not a product category, and the post is precise about it: vector search is not dying, the vector-primary storage layout is. The distinction matters for anyone buying infrastructure. A vector database is still the right abstraction for similarity search; the argument is that its storage engine should treat embeddings as one index among several, keyed off a document primary, so filtered search, full-text, and vector queries compose without duplicating data or dragging indexes across cluster rebalances. That is the architecture the post-SQL-generation of search engines has been converging on, and turbopuffer running it at 100B-vector scale makes the strongest case yet that the convergence is not a compromise.

The Honest Caveats Are the Credibility

The post ships with its own limits stated plainly: v3 has passed 100% of CI but is at “day zero of perf grinding,” with no performance benchmarks yet and public ones promised before production rollout. The team flags the real risk, that ANN performance could regress, since the vector-primary layout works “really, really well” for pure vector workloads on object storage. Even the validation step is disclosed: the Opus-5 control arm swap in their harness was statistically indistinguishable at 99% confidence, an admission that their own instrumentation has limits. Engineering blogs that publish rearchitectures at this stage of honesty, before the benchmark victory lap, are rarer than the rearchitectures themselves.

What to Watch

First, the v3 benchmarks, which will show whether document-primary layout costs ANN speed and by how much; that delta is the price of composability, and buyers will decide with it. Second, whether the pattern spreads: if document-primary with vectors-as-secondary becomes the default across retrieval infrastructure, the standalone vector database category shrinks into a feature, which is what the title is really predicting. Third, the full-SQL roadmap: v3 “lays the foundation” for more SQL in turbopuffer, and a retrieval engine that speaks SQL natively over billion-vector indexes would collapse yet another layer of the modern RAG stack. The RAG infrastructure market just got its clearest signal yet that the 2023 stack diagram is obsolete.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading