tpuf is a small, educational clone of turbopuffer: a vector +
full-text search engine built on one bet β object storage is the source of truth, and everything
else exists to hide its latency.
It is a Go core-engine library plus a CLI (create | upsert | index | query | info | branch) that does
centroid/IVF vector search and BM25 full-text search over MinIO (S3-compatible) running in
Docker. The goal is a faithful core in readable code β not production scale. When a choice is between
clever and clear, this codebase chooses clear.
This README summarizes the design. The full, sourced knowledge base lives in
docs/β start atdocs/README.md. Where turbopuffer does not publicly confirm a number, the docs flag it explicitly; this README does the same and never asserts an unconfirmed internal as fact.
Incumbent search engines keep data on RAM plus triply-replicated SSD, inheriting architectures built for low-latency transactional updates. turbopuffer's observation: search is not a transactional database β its write workload looks like a data warehouse (high throughput, no transactions, relaxed write latency), with the one extra requirement that reads come back fast (turbopuffer targets under ~100 ms). So it pairs warehouse-style storage (S3/GCS, ~$0.02/GB) with search-style read latency (a smart index plus caching).
Three ideas make that work, and tpuf implements all three:
- Object storage is the source of truth. A namespace is just a key prefix
(
{bucket}/{namespace}/...) holding a write-ahead log and a built index. Cheap, durable, and (since S3 went strongly consistent in Dec 2020 and added conditional writes) coordinatable. - Coordination is CAS on a single JSON file β no Raft, no Kafka, no Zookeeper. A
manifest.jsonis updated with a conditionalIf-Match: <etag>PUT. If the ETag still matches, the write wins; if not, the store returns HTTP 412 Precondition Failed and the writer reloads and retries. That is exactlyUPDATE ... WHERE version = N, delegated to object storage. (MinIO supports bothIf-MatchandIf-None-Match, so the clone implements the real CAS model faithfully.) - Unindexed data is still searchable. Indexing is asynchronous. Until a write is folded into the index, queries find it by exhaustively scanning the unindexed WAL tail and unioning the results.
Sources for the above: the turbopuffer founding blog and architecture/concepts pages, distilled in
docs/01-architecture.md.
A namespace is fully isolated under its own prefix:
{bucket}/{namespace}/
βββ manifest.json # the CAS-coordinated source-of-truth head (version, dim, metric,
β # WALSeq, IndexedUpTo, IndexEpoch, DocCount)
βββ wal/00000000000000000000.json # committed write segments (the LSM log), 20-digit zero-padded seq
βββ wal/00000000000000000001.json
βββ index/v{epoch}/ # one immutable, write-once index epoch
βββ centroids.json # centroid vectors + cluster sizes
βββ cluster-0.json # vectors (+ RaBitQ-lite codes) + attrs for cluster 0
βββ bm25.json # inverted index for full-text
βββ docs.json # id -> attributes (filtering + return payloads)
The manifest.json is the head: its ETag is the CAS token, and it points at the live index epoch and
the WAL high-water mark.
The repo follows the official go.dev "single command + supporting packages" shape (see
docs/07): a CLI in cmd/, the engine in internal/.
cmd/tpuf/main.go CLI: create | upsert | index | query | info | branch
(query: --vector and/or --bm25; both β hybrid RRF)
cmd/tpuf-bench/main.go latency benchmark: p50..p99.9, single/multi-tenant/cold-vs-hot +
per-feature demos (--hybrid/--group-commit/--filter-plan/--rabitq/--nvme-dir)
cmd/tpuf-node/main.go stateless HTTP query server (for the load-balancer demo)
cmd/tpuf-broker, tpuf-indexer async-indexing daemons coordinating via CAS-guarded queue.json
internal/bench/bench.go latency Recorder + nearest-rank percentile summary + table renderer
internal/storage/
storage.go ObjectStore interface (Get/PutCAS/PutIfAbsent/Put/List) + sentinel errors
s3.go S3(MinIO) impl with CAS (If-Match) β the only file importing the AWS SDK
memory.go in-memory ObjectStore with real CAS β infra-free tests & a no-Docker demo mode
internal/cache/
cache.go DRAM-tier cache over object storage (immutable index objects only)
nvme.go NVMe FIFO ring-buffer middle tier (the 3-tier DRAM/NVMe/S3 story)
internal/engine/
types.go Document, WALSegment, Manifest, RankBy, Filter, QueryResult, on-disk index shapes
manifest.go load + CAS-save the manifest; branch fork
wal.go append / list / read WAL segments; materialize live docs (last-writer-wins)
vector.go cosine/euclidean distances, k-means, true-RaBitQ rotation+estimator, centroid tree
bm25.go tokenizer, inverted index, BM25 scoring
indexer.go build centroid + BM25 + bitmaps + docs index from the WAL, then CAS the epoch swap
query.go planner: vector/BM25/hybrid-RRF + bitmap filter plan + WAL-tail scan + merge
namespace.go Namespace handle: Create, Upsert, Index, Query, Info, Branch
bitmap.go branch.go commit.go queue.go lire.go the extension implementations
deploy/ nginx LB demo + broker/indexer (indexer profile, queue-demo.sh)
benchmarks/ run.sh (canned bench + feature-demo configs) + RESULTS.txt (saved runs)
- Write (
upsert) β durable before return. Append a newwal/{seq}.jsonsegment (write-once, viaIf-None-Match), then CAS the manifest to bumpWALSeq. The call returns only after the store acknowledges; a successful return means the data is durably on object storage. - Index (
index) β build a fresh epoch. SnapshotWALSeqat the start, materialize the live docs, run k-means (K β βN clusters) and BM25, write everyindex/v{epoch}/*object under the new prefix, then publish with a single manifest CAS that flipsIndexEpochand setsIndexedUpToto the start snapshot. Until that one CAS lands, queries keep serving the old epoch. - Query (
query) β one endpoint, mode chosen byRankBy: a vector triggers ANN (probe the topnProbeclusters, RaBitQ-lite prefilter, exact rerank); text triggers BM25 over the inverted index. Either way the query then exhaustively scans the unindexed WAL tail[IndexedUpTo, WALSeq), applies last-writer-wins (newer overwrites the indexed copy), subtracts tombstoned deletes, applies metadata filters, sorts, and returns top-K. With no index yet (IndexEpoch == 0) it serves purely from the tail β proof that unindexed data is searchable.
You need Docker (for MinIO) and Go 1.26.
# 1. Boot MinIO (S3 API :9000, console :9001) and auto-create the `tpuf` bucket.
docker compose up -d
# 2. Load the S3 credentials/endpoint into your shell.
set -a; source .env.example; set +a
# 3. Build and run the engine's unit tests (fast; no infra needed thanks to the in-memory store).
go mod tidy && go build ./... && go test ./internal/engine/...
# 4. Create a namespace: 4-dim cosine vectors, full-text on the `body` attribute.
go run ./cmd/tpuf create demo --dim 4 --metric cosine --text-field body
# 5. Upsert the sample docs (4-dim vectors + a `body` text field + a `lang` attribute).
go run ./cmd/tpuf upsert demo --file examples/sample.json
# 6. Query BEFORE indexing β answered entirely from the WAL tail (the headline demo).
go run ./cmd/tpuf query demo --vector "0.1,0.2,0.3,0.4" --top-k 3
go run ./cmd/tpuf query demo --bm25 "quick walrus" --top-k 3
# 7. Build and publish an index epoch via a single manifest CAS.
go run ./cmd/tpuf index demo
# 8. Query AFTER indexing β the indexed path plus the (now empty) WAL tail.
go run ./cmd/tpuf query demo --vector "0.1,0.2,0.3,0.4" --n-probe 3
go run ./cmd/tpuf query demo --bm25 "walrus" --filter '{"op":"eq","field":"lang","value":"en"}'
# 9. Inspect the manifest.
go run ./cmd/tpuf info demo
# 10. Open the MinIO console to see every object as it is written.
# http://localhost:9001 (login with MINIO_ROOT_USER / MINIO_ROOT_PASSWORD)
# You'll see demo/manifest.json, demo/wal/..., and demo/index/v1/...The two proofs to watch for: step 6 returning results with no index present shows unindexed WAL data is searchable; the MinIO console showing every object shows object storage is the source of truth.
gofmt -w . && go vet ./... # format + vet
go test ./... # full unit suite β fast, no infra
go test ./... -race # exercises the CAS / concurrent-upsert paths
go test ./internal/storage -tags=integration # real-MinIO 412 contract test (needs Docker)tpuf-bench builds a fresh namespace, upserts synthetic docs in batches, then times vector and BM25
queries before indexing (each query scans the unindexed WAL tail) and after (served from the
indexed epoch with the DRAM cache warm). The gap between those rows is the whole thesis made visible.
go run ./cmd/tpuf-bench --backend memory --docs 2000 --batch 100 --queries 300 # pure engine cost, no infra
set -a; source .env.example; set +a # then, against real MinIO:
go run ./cmd/tpuf-bench --backend s3 --docs 500 --batch 50 --queries 50 # real object-storage latencyA representative MinIO run: the vector tail scan lands around p50 15ms / p99 19ms (a manifest read plus one object read per WAL segment, every query β the WAL is never cached), while the indexed path is ~p50 3.8ms / p99 4.6ms (only the manifest GET hits MinIO; index objects come from the DRAM cache). Latency is dominated by object-storage round-trips, which is exactly what the index and cache exist to hide.
Multi-tenant mode (--namespaces N --concurrency C) is the realistic regime: N tenants are each
created + indexed, then C worker goroutines issue queries spread across them, sharing one DRAM cache.
Pair it with --cache-objects M β a bounded LRU cache β set below the resident working set to watch
cold-start misses recur under tenant churn (the pressure a finite DRAM tier, and the real product's
omitted NVMe tier, must absorb):
go run ./cmd/tpuf-bench --backend s3 --namespaces 12 --concurrency 12 \
--dim 256 --docs 1000 --queries 6000 --cache-objects 80 # bounded: forces eviction
go run ./cmd/tpuf-bench --backend s3 --namespaces 12 --concurrency 12 \
--dim 256 --docs 1000 --queries 6000 --cache-objects 0 # unbounded: every tenant stays residentOn a terminal, multi-tenant runs show a live Bubble Tea dashboard (per-phase progress bars, throughput, cache hit-rate, and ETA); piped or non-TTY runs fall back to plain periodic progress lines. It reports aggregate p50β¦p99.9 under load, achieved wall-clock throughput, and the shared cache's hit/miss/eviction counts. Capping the cache below the working set drops the vector hit rate sharply (e.g. ~99% β ~53%) because each tenant's vector index is many objects (centroids + cluster files), while BM25 β two objects per tenant β stays hot.
There is also a canned runner and saved reference numbers:
benchmarks/run.sh # memory backend, single-tenant (no infra)
benchmarks/run.sh --backend s3 --mode all # single + multi-tenant + cold-vs-hot against MinIOFull output and methodology in benchmarks/RESULTS.txt. Single-tenant and
cold-vs-hot rows re-measured 2026-06-18 (post-extensions). The headline findings:
| What | Measurement | Takeaway |
|---|---|---|
| Index vs. WAL-tail scan (single tenant) | vector query p50 15ms β 3.8ms, p99 19ms β 4.6ms | indexing + DRAM cache cut tail latency ~4Γ |
| Cold vs. hot, same query | vector p50 17.2ms (cold) β 11.0ms (hot), 1.6Γ; cold pass = 480 cache misses, hot pass = 0 | the cache's win is the avoided object-storage round-trip |
| Cache pressure (12 tenants, dim 256) | bounded 80-obj cache 53.8% hot, p50 38.8ms, 297 qps β unbounded 99.7% hot, p50 21.9ms, 520 qps | a finite DRAM tier degrades gracefully; the manifest read is the latency floor |
| Scaling (consistent hashing) | add a 4th node β 40/50 namespaces stay put, the 10 that move all go to the new node | scaling doesn't cold-flush existing caches (plain hash % N would move ~37/50) |
All numbers are single-node loopback MinIO and swing with machine load; read the ratios (indexed βͺ tail, hot βͺ cold), not the absolute milliseconds. On real cloud object storage the per-query manifest GET alone is tens of ms β the floor turbopuffer's NVMe tier and <100 ms target exist to manage.
Caveat the benchmark surfaces honestly: every query does one uncached manifest GET (correctness rule 2), so even a 100%-cache-hit query has a ~object-storage-round-trip floor. Real turbopuffer optimizes that away; the clone leaves it visible.
An optional demo of turbopuffer's routing tier β nginx consistent-hashing the namespace to one of
several identical Go query nodes (so a tenant's cache stays warm on one node, while any node can serve
any namespace because state is on S3). Full walkthrough in deploy/README.md:
docker compose --profile lb up -d --build # 3 nodes + nginx, reusing MinIO (LB on :8088)
deploy/demo.sh # show hash(namespace) -> node + per-node cache warmth
deploy/remap-demo.sh # add a node, watch only ~1/N namespaces remapThe real turbopuffer is written in Rust, chosen for predictable low-latency performance, tight memory control, and no GC pauses on the query hot path β the properties you want when you are squeezing ranged reads out of object storage at scale.
This clone is written in readable Go on purpose. The trade is deliberate: Go's standard library,
simple concurrency, and low ceremony make the architecturally interesting mechanisms β the CAS retry
loop, the WAL-tail scan, the two-step centroid retrieval, the binary-prefilter rerank β easy to read end
to end. We are not chasing turbopuffer's latency or throughput numbers; we are reproducing its
shape so you can trace every code path. Concretely, the clone leans on Go idioms that aid clarity over
raw speed: the engine talks only to a small ObjectStore interface (so the whole thing is testable
against an in-memory store with no Docker), errors are wrapped and matched with errors.Is/errors.As,
and the 412/404 contract is the single load-bearing concurrency primitive rather than a custom runtime.
Object-storage-first persistence; namespaces as isolated prefixes; durable-before-return WAL writes;
CAS-on-a-JSON-file coordination with real If-Match/412 retries; async indexing published via an atomic
epoch swap; exhaustive WAL-tail scan so fresh writes are searchable; a centroid (IVF) vector index with
two-step retrieval; a RaBitQ-style binary prefilter then exact rerank; hand-written BM25 (k1=1.2,
b=0.75); a single Query endpoint whose mode is chosen by RankBy; and strong consistency by default
(the query routes through the manifest, so new writes are always visible).
docs/05-clone-mapping.md listed these as deliberate non-goals; as of
2026-06-18 they are built under docs/extensions/ (design rationale) +
docs/extensions/IMPLEMENTATION-HANDOFF.md (build map),
each with co-located -race tests and a cmd/tpuf-bench or CLI/deploy/ demo:
- Group commit β opt-in per-namespace committer coalesces concurrent upserts into one WAL PUT
(
internal/engine/commit.go;tpuf-bench --group-commitshows ~8Γ fewer PUTs). - Live broker + indexer processes /
queue.jsonβ async-indexing daemons claim jobs via the same CAS shape as the manifest (cmd/tpuf-broker,cmd/tpuf-indexer; composeindexerprofile). - SPFresh LIRE incremental updates β Phase 1 (Option A: incremental rebuild β split/merge/
reassign + per-vector version β behind the unchanged single-CAS epoch) in
internal/engine/lire.go; Option B/C deferred. - Hierarchical centroid tree β recursive k-means levels in
centroids.json(vector.go); measured fan-out honestly shows it's pedagogical at this N. - Bitmap / attribute indexes + filter-first vs search-first planner β
index/v{epoch}/bitmaps.json(internal/engine/bitmap.go); same answers as the per-candidate path, less I/O when selective. - NVMe ring-buffer cache tier β a FIFO disk tier under the DRAM map (
internal/cache/nvme.go), making the real 3-tier DRAM/NVMe/S3 story (tpuf-bench --nvme-dir). - True RaBitQ rotation + unbiased estimator (
vector.go), hybrid RRF fusion (tpuf query --vector --bm25), and copy-on-write branches (tpuf branch,internal/engine/branch.go). (The consistent-hashing load-balancer demo underdeploy/remains too.) Still documented-not-built: CMEK, multi-region, LIRE Option B, weighted score fusion, branch GC.
These additions all preserve the behavior that makes the architecture interesting (and the 5 CAS correctness rules); the scale machinery is implemented at clone scale to show its shape, not to chase billions-of-vectors performance.
The engine, CLI, benchmark tool, load-balancer demo, all 8 docs/extensions features, and SPFresh
LIRE Phase 1 are implemented and tested: go test ./... -race is clean (165 tests, no
infrastructure), and the real-MinIO contract test passes under -tags=integration. The full end-to-end
recipe in docs/06 β plus hybrid queries, attribute filters,
and the copy-on-write branch flow β is verified against MinIO. The research and design remain sourced in
docs/; the extension build map is
docs/extensions/IMPLEMENTATION-HANDOFF.md.
The engine's single external dependency is aws-sdk-go-v2/service/s3. Everything conceptually
interesting β vector math, k-means, the RaBitQ-lite codes, BM25, and the CAS loop β is hand-written
standard-library Go, because reaching for a library to implement a core concept would defeat the point of
the clone. The one exception is the optional benchmark CLI (cmd/tpuf-bench), which uses
charmbracelet/bubbletea + bubbles + lipgloss for its live progress TUI β dev tooling that sits
outside the engine, so the engine's single-dependency invariant holds.