Skip to content

< CurriculumRetrieval · 32 of 46 ·76 · Level 02, Search

Filter by Metadata Before You Rank

medium · implement · Embeddings & Retrieval

Asked what was the travel expense limit in 2023?, the assistant gives the 2025 figure. The three years' memos are the same sentence with a different number, so to an embedding model they are the same document. Exact conditions like a year belong in a filter on the memos' metadata; similarity ranks whatever survives.

Each memo is {"id", "dept", "year", "text"}; dept is one of finance, hr, it, legal. Implement two functions.

extract_filters(question) returns only the conditions the question states:

  • "year": an int, when the question names a year.
  • "dept": one of the four codes, when the question names a department. People write human resources, the lawyers, tech support, accounts, not the codes.
  • Neither named: {}. Do not guess a department from the topic; travel expenses is a topic, not a department.

llm.ask(prompt) (the chat model), jev (the decision model) and plain Python are in scope. Think about which part needs a model at all.

search(question, memos, k=3) returns the ids of the k best memos, best first:

  1. Get the filters and keep only the memos matching all of them.
  2. Rank the survivors by similarity to the question with embedder.embed(texts) and cosine(a, b); return the top k.
  3. Nothing survives: return [].

The catch: the model is too helpful. Asked which department a question names, it will name one when there is none, because working from home sounds like HR. Have it quote the words it relies on, and check the quote is really in the question.

Filter first, then rank: ranking first and then dropping the wrong years leaves one result or none, and the tests check for exactly that. They also count calls across all tools: at most three per search.