Skip to content

< all problems70 · Level 01, LLM APIs

Put the Answer Where the Model Will Read It

medium · implement · LLM Fundamentals

A staff handbook assistant answers from eight policy documents. Somebody built its prompt by pasting in all eight, in filing order, with the question on top. It answers a question about holiday with the parking policy's numbers about one time in five.

Two things are wrong, and neither is the model. Most of that prompt is irrelevant to any one question, and the part that matters sits wherever the filing order happened to put it. A model does not read a long prompt evenly: it attends most to the beginning and the end, and least to the middle.

Implement build_prompt(question, docs, k=3), returning one string. docs is a list of {"id": ..., "text": ...}.

The prompt has three parts, separated by blank lines:

  1. INSTRUCTION (provided), first.
  2. The k documents most relevant to the question, each written as <doc id="ID">TEXT</doc>, ordered so that the most relevant is last: nearest the question, where it will be read. The other documents are left out.
  3. The line Question: followed by the question, last.

How you decide what is relevant is up to you. In scope: embedder.embed(texts) with cosine(a, b), the real embedding model; jev.ask(state, questions), the decision model; llm.ask(prompt), the chat model; or plain Python. Staff ask I want to spend a month at my family's place abroad and keep working; the policy says staff may work from another country for up to 20 working days. Hardly a word in common.

The tests read the prompt you built: the instruction opens it, the question closes it, there are exactly k documents, and over six questions the document that answers each must be the last one for at least five. They also count calls, across all the tools: at most two per question. Then a real model is given your prompt and has to get the answer right.