Answer Over a Context That Does Not Fit
A model that answers reliably over two thousand tokens starts missing facts buried in two hundred thousand. A recursive language model never shows the root call the whole context: it splits it, asks a model about each part, and combines the partial answers.
Implement recursive_answer(llm, question, context, max_depth=2, chunk_chars=1200):
- Base case. If the context fits within
chunk_chars, ormax_depthhas run out, answer it directly in one call. - Split the context into chunks of at most
chunk_chars, on paragraph boundaries so sentences survive. - Recurse on each chunk with
max_depth - 1, collecting partial answers. - Combine the partials in one final call to produce the answer.
The catch: the base case must trigger when depth is exhausted regardless of size, or a long context recurses until something breaks. And most chunks do not contain the answer, so tell the sub-call to say so briefly rather than guess, or you combine four confident wrong answers with one right one.
You are scored on model_calls as well as correctness. Recursion multiplies calls fast: depth 2 over 4 chunks is already 6 calls.