Skip to content

< CurriculumContext engineering · 18 of 46 ·70 · Level 01, LLM APIs

Put the Answer Where the Model Will Read It

medium · implement · LLM Fundamentals

A staff handbook assistant answers from eight policy documents, all pasted into the prompt in filing order with the question on top. One time in five it answers a holiday question with the parking policy's numbers: most of the prompt is irrelevant, and a model attends least to the middle of a long prompt.

Implement build_prompt(question, docs, k=3), returning one string. docs is a list of {"id": ..., "text": ...}. The prompt has three parts, separated by blank lines:

  1. INSTRUCTION (provided), first.
  2. The k documents most relevant to the question, each written as <doc id="ID">TEXT</doc>, with the most relevant last, nearest the question. The other documents are left out.
  3. The line Question: followed by the question, last.

How you decide relevance is up to you. In scope: embedder.embed(texts) with cosine(a, b), the real embedding model; jev.ask(state, questions), the decision model; llm.ask(prompt), the chat model; or plain Python. Beware that a month at my family's place abroad and keep working and work from another country for up to 20 working days share hardly a word.

The tests read the prompt: the instruction opens it, the question closes it, there are exactly k documents, and over six questions the answering document must be last for at least five. They count calls across all the tools: at most two per question. Then a real model is given your prompt and must get the answer right.