Skip to content
AnalysisDT-2026-0192

Chunking is still the whole game

Teams swap vector databases hoping for better recall. The gains almost always came from how the text was split.

9 minBig Data & Vector DBs

A retrieval system has three parts that can be wrong: how documents are split, how they are embedded, and how results are ranked. Of the three, splitting receives the least attention and accounts for most of the variance we measure.

This is partly because chunking is unglamorous and partly because it is domain-specific. A fixed 512-token window is defensible for prose and indefensible for a table, a contract clause or a stack trace.

What actually moved recall

  • Splitting on document structure rather than token count — headings, clauses, table rows.
  • Keeping a parent reference so a matched fragment can be expanded to its section at answer time.
  • Storing a short generated summary alongside each chunk and embedding both.
  • Reranking with a cross-encoder over the top fifty rather than trusting the first five.

Where the database choice matters

It matters for operations, not for quality. Filtering semantics, index rebuild behaviour under write load, and whether hybrid search is a first-class feature or a bolt-on — these differ sharply between systems and will shape your on-call rota.

We migrated three times before accepting that our recall problem had nothing to do with the store.
Data platform lead, legal technology company

Evaluate the retriever on its own, with a fixed set of questions and known-good passages, before changing anything underneath it. Most teams skip this step and then attribute the result to whichever component they replaced last.

What to watch

Structure-aware splitting is becoming a product feature rather than a pipeline you write. Whether that commoditises the advantage or just raises the floor is the open question for the next two quarters.

Read next

Across the network

Desks that share a zone with this one on the BITBRIEF coverage map.

Terms defined