← Back

RAG Chunking: Preserve Meaning Before Tuning Size

Choose a RAG chunking strategy around document structure, evidence boundaries, and retrieval tests. Learn when overlap helps and when it creates duplicate noise.

A RAG chunking strategy should preserve the information needed to answer a question. Token counts matter, but they are a constraint rather than a definition of meaning. A beautifully sized chunk that separates a rule from its exception is a bad chunk for that question.

Imagine a deployment guide with three consecutive sections: prerequisites, migration commands, and rollback conditions. A fixed splitter can place the commands in one chunk and the warning about an incompatible schema in another. Retrieval may find the commands perfectly while giving the assistant an incomplete procedure.

Identify the evidence unit

Start by asking what a useful answer needs to cite. For an API reference, that might be one endpoint with its parameters and response semantics. For an incident report, it might be a timeline event with the surrounding explanation. For a policy, it may be a rule plus its exclusions.

My default design recommendation is to preserve headings and paragraph boundaries first, then split sections that exceed the embedding or generation budget. Treat tables, code blocks, and numbered procedures as structures with dependencies. Repeating a table's column headers in a split segment is more useful than copying an arbitrary number of preceding characters.

Anthropic's contextual retrieval article describes adding short, chunk-specific context before indexing. Its central motivation is that isolated passages can lose the identity of the document or subject they describe. That is a useful design option, not evidence that generated context is necessary for every collection.

Keep retrieval text and source evidence distinct

Store the original passage separately from any text added to improve retrieval. A generated prefix may help match a question, but it is not itself an authoritative sentence from the document. If the prefix invents a product version, you need to locate and remove that error without rewriting the source.

An illustrative chunk record might contain:

{
  "document_id": "deployment-guide",
  "version": "7",
  "heading_path": ["Migrations", "Rollback"],
  "source_start": 1420,
  "source_end": 2190,
  "retrieval_prefix": "Rollback requirements for schema v7",
  "source_text": "...original passage..."
}

Offsets here are illustrative. In an implementation, define whether they count bytes, characters, or tokens, and preserve the exact source snapshot. A citation offset into a document that has since changed is not a stable reference.

The generation step can receive the heading path and original passage while treating generated retrieval metadata as secondary. Review source excerpts, not just search previews, when deciding whether a chunk supports an answer.

Use overlap for a specific reason

Overlap helps when a sentence near a boundary depends on nearby text. It also increases storage and can cause several nearly identical candidates to occupy the final context. Those costs are not automatically bad; they need a purpose.

Consider a 900-word policy section with a short exception at the end. Copying the last paragraph into the next chunk may preserve the exception locally. Copying 30 percent of every chunk across an entire corpus is a much broader decision. Test it against questions where boundaries actually caused misses.

Deduplicate at context assembly, using document identity and source ranges. Two overlapping chunks may both score highly because they repeat the same answer. Their rank should not displace a different passage that supplies a necessary limitation.

When combining adjacent passages, preserve provenance. A merged excerpt should still let the reader navigate to its original document and location. Silent concatenation makes later debugging unnecessarily difficult.

Test several plausible strategies

Compare a small number of deliberate candidates: section-based splitting with a size cap; smaller child chunks retrieved with a larger parent section; and a fixed-size baseline. Keep the source documents, embedding model, query set, and ranking settings constant.

Use questions that require exact identifiers, cross-paragraph context, table interpretation, and exceptions. Measure whether the required evidence reaches the final prompt, not merely whether the search results contain the right document somewhere.

Sentence Transformers' retrieve-and-rerank guide separates candidate discovery from final relevance scoring. That distinction is useful here: a chunking change can help the first stage while burdening the second with duplicates. Inspect both stages before calling it an improvement.

Do not compare strategies solely by average answer length or similarity to a reference answer. A concise answer can still omit the one caveat that matters.

Handle long context deliberately

Returning a whole parent section is attractive when child chunks lose meaning. It can also consume the answer budget with unrelated details. Place a cap on expanded context and prefer sections whose additional text explains the retrieved evidence.

Lost in the Middle documents sensitivity to the position of relevant information in the evaluated long-context tasks. Its findings should motivate position-sensitive testing, rather than a universal claim about every current model. Try moving the decisive passage within your assembled context and examine whether the answer changes.

A chunking system should therefore record both what it retrieved and what it ultimately sent. If the rollback warning was retrieved but trimmed during prompt assembly, changing the splitter is solving the wrong problem.

Make the ingestion contract explicit

Document how headings, lists, tables, source versions, and access rules survive ingestion. Add a few human-inspected fixtures for your hardest document shapes. Re-run those fixtures whenever the parser or splitter changes.

Use the RAG evaluation workflow to compare chunking changes against a fixed evidence set. The DB-GPT overview places that retrieval work inside a broader data assistant.

Choose the smallest chunking design that preserves your evidence units and passes your cases. A number such as 512 tokens is an experiment setting. It is not a document-understanding strategy.