Back
turbopuffer retires the vector-primary index in v3 storage rework
SiTech AI Team3 min read

turbopuffer retires the vector-primary index in v3 storage rework

In a blog post, turbopuffer says its upcoming v3 storage architecture stops keying documents and indexes on ANN vector addresses, demoting the vector index to a secondary role, and reports 100% of CI passes on v3.

Serverless search engine turbopuffer is changing how it stores and indexes data. In a blog post titled "RIP, vector database", published on September 30, engineer Dan Harrison described turbopuffer v3: a storage architecture in which documents and indexes are no longer keyed on an ANN (approximate nearest neighbor) vector address, the vector index steps down to a secondary role, and a new primary index takes its place.

The company says the change will speed up every kind of search, including text, regex and vector, and lays the foundation for running many more SQL queries on turbopuffer.

From v1 to v2

turbopuffer launched as a serverless vector database whose documents consisted of an ID and a vector, with object storage as the source of truth and tiered NVMe and memory caches for performance. Vectors were organized in a hierarchical clustering tree: SPANN first, later SPFresh for incremental indexing. Early customers including Cursor and Notion validated those trade-offs. The second generation added attribute filtering and BM25 full-text search, then regex search, fuzzy matching, sparse vector search and aggregations; it also powers non-search workloads such as Linear's syncing engine. The query engine grew, but the storage layout stayed vector-primary, constraining query plans like GROUP BY and aggregations.

Three costs of a vector primary index

The post names three problems. Storage amplification: a document's full contents live under its vector address, so multi-vector representations such as document nesting or late interaction duplicate the same data for every vector. Write amplification: SPFresh rebalances clusters on every insert, update or delete, and because attributes and inverted indexes are keyed the same way, updating one vector can move hundreds of attributes and their indexes. Limited vectorization: modern engines process blocks of values (DuckDB 2,048 rows, ClickHouse about 65k, Lucene 256-doc blocks), while turbopuffer's ANN clusters work best at 100-200 documents, so plans that prefer larger blocks stay capped at cluster size.

The company cites its own full-text search: moving postings from cluster-based blocks, where the median held about 1.5 postings, into fixed blocks of about 256 made the index 10x smaller and queries up to 20x faster.

What comes next

v3 is not in production yet. Earlier in September, 100% of CI runs passed on v3, closing the correctness phase; the team now intends to make it fast, publishing benchmarks at turbopuffer.com/v3 over the coming weeks as it works toward, and beyond, performance parity before production. turbopuffer says it hosts more than 1 trillion documents, handles over 10 million writes per second and serves more than 25,000 queries per second; the previous architecture reached single indexes of more than 100 billion vectors at 200 ms p99 reads and over 1,000 queries per second.

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.