Put the Answer Where the Model Will Read It
A staff handbook assistant answers from eight policy documents, all pasted into the prompt in filing order with the question on top. One time in five it answers a holiday question with the parking policy's numbers: most of the prompt is irrelevant, and a model attends least to the middle of a long prompt.
Implement build_prompt(question, docs, k=3), returning one string. docs is a list of {"id": ..., "text": ...}. The prompt has three parts, separated by blank lines:
INSTRUCTION(provided), first.- The
kdocuments most relevant to the question, each written as<doc id="ID">TEXT</doc>, with the most relevant last, nearest the question. The other documents are left out. - The line
Question:followed by the question, last.
How you decide relevance is up to you. In scope: embedder.embed(texts) with cosine(a, b), the real embedding model; jev.ask(state, questions), the decision model; llm.ask(prompt), the chat model; or plain Python. Beware that a month at my family's place abroad and keep working and work from another country for up to 20 working days share hardly a word.
The tests read the prompt: the instruction opens it, the question closes it, there are exactly k documents, and over six questions the answering document must be last for at least five. They count calls across all the tools: at most two per question. Then a real model is given your prompt and must get the answer right.