Build a Document Chunker
Implement chunk_document(text, chunk_size, overlap), splitting text into a list of word chunks. Assume chunk_size > 0 and 0 <= overlap < chunk_size.
- Each chunk holds at most
chunk_sizewords. - Each chunk after the first starts
overlapwords before the previous one ended. - Chunks are strings, words rejoined with single spaces.
- Never an empty chunk, and never a chunk whose words are all already covered by the previous one.
- Empty or whitespace-only input returns
[].
The catch: rule 4. A naive range(0, len(words), step) emits a redundant tail chunk whenever the text divides unevenly.