Skip to content

< CurriculumContext engineering · 24 of 46 ·28 · Level 08, Production

Answer Over a Context That Does Not Fit

hard · implement · RLM

A model that answers reliably over two thousand tokens starts missing facts buried in two hundred thousand. A recursive language model never shows the root call the whole context: it splits it, asks a model about each part, and combines the partial answers.

Implement recursive_answer(llm, question, context, max_depth=2, chunk_chars=1200):

  1. Base case. If the context fits within chunk_chars, or max_depth has run out, answer it directly in one call.
  2. Split the context into chunks of at most chunk_chars, on paragraph boundaries so sentences survive.
  3. Recurse on each chunk with max_depth - 1, collecting partial answers.
  4. Combine the partials in one final call to produce the answer.

The catch: the base case must trigger when depth is exhausted regardless of size, or a long context recurses until something breaks. And most chunks do not contain the answer, so tell the sub-call to say so briefly rather than guess, or you combine four confident wrong answers with one right one.

You are scored on model_calls as well as correctness. Recursion multiplies calls fast: depth 2 over 4 chunks is already 6 calls.