Rerank With a Model That Understands the Question
A customer asks does the warranty cover water damage? and the retriever returns six passages, every one about the warranty or about water. The top one, a common question is whether the warranty covers water damage, answers nothing; the one that answers, damage from liquids is not covered, is fifth. A passage that repeats the question is as similar to it as text can be, and being about the question is not answering it.
Implement rerank(question, passages, k=3). passages is a list of {"id": ..., "text": ...} in the retriever's order. Return the ids of the k passages that best answer the question, best first.
How you judge is up to you. In scope:
jev.ask(state, questions): the decision model. It answers typed questions with numbers, and one call can carry a question for every passage.llm.ask(prompt): the chat model. Show it the question and the numbered passages and ask for an order.embedder.embed(texts)withcosine: the embedding model. Fast and cheap. Find out whether nearness is enough here.
The tests check the result, not the method: over six questions, each with six passages on its topic, the passage that answers must come first for at least five. They count calls across all the tools: at most two per question.