Skip to content

< CurriculumRAG and grounding · 46 of 46 ·81 · Level 03, RAG

Answer From a Repo That Does Not Fit

hard · implement · RAG

Asked what is the default retry limit?, the assistant pastes the whole repository into the prompt. The repository is twenty-three files, the model's context window is 900 tokens, and the call fails before the model reads a line. Put the right 900 tokens in front of the model, not all of them.

REPO (provided) is the repository, a dict of path to file contents. model (provided) is the chat model behind a window of model.window tokens, measured with count_tokens (provided): a prompt over the limit raises ContextWindowExceeded instead of being sent. embedder.embed(texts) and cosine(a, b) are the real embedding model. Implement three functions.

chunk_repo(repo) returns a list of {"path", "text"} chunks, each at most CHUNK_TOKENS (provided, 160) tokens and from one file, with every line of every file in exactly one chunk, in order. A file that fits is one chunk; split a longer file on line boundaries.

build_index(chunks) embeds every chunk in one call and returns an index that keeps each chunk with its vector.

answer(question) returns {"answer": ..., "files": [...]}:

  1. Embed the question (one call) and rank the chunks by cosine similarity.
  2. Pack the prompt: an instruction, then chunks in order of relevance, each labelled with its path, stopping before the next chunk would push the whole prompt over model.window. Leave room for the question. Count with count_tokens; do not guess.
  3. Ask the model to answer from those chunks in one sentence. Return the answer and the paths of the chunks you sent, each once, most relevant first.

Build the index once, not once per question: fill INDEX at module level, or build it lazily on the first call.

Several files mention retries, timeouts, caches and limits. The tests check that the answer carries the value asked for, that its file was among those sent, and that no call exceeded the window (model.largest_prompt records the biggest prompt you sent).