Choose chunk boundaries by whether a retrieved chunk can answer a real query with enough context.
Chunk size is a retrieval design decision, not a tokenizer chore. Specificity Shorter chunks can match narrow questions because less unrelated text dilutes the vector. They are useful for definitions, FAQ answers, ticket snippets, and product attributes. Context Longer chunks preserve conditions, exceptions, examples, and table meaning. They help when the answer depends on surrounding language, but they can also retrieve for the wrong sentence inside the chunk. Structure Headings, tables, code blocks, legal sections, and policy exceptions carry meaning. Structure-aware chunking usually beats blind fixed-size splitting when documents have real organization. Expansion When a precise child chunk is not…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in