Skip to content

< all problems68 · Level 01, LLM APIs

Pick the Examples That Fit the Question

medium · implement · LLM Fundamentals

A support desk sorts incoming tickets into queues with internal codes: SHIP-1, PAY-4, AUTH-3 and so on. A model does the sorting, and nobody has ever told it what the codes mean. It learns them the only way it can: from worked examples placed in the prompt, a ticket and then its code. That is few-shot prompting.

There are twelve worked examples in POOL and room in the prompt for two. Put in two about passwords when the ticket is about a late parcel, and the model has never seen SHIP-1; it will pick one of the codes it was shown, confidently, and the ticket goes to the wrong team. Which examples go in decides the answer.

Implement two functions.

pick_examples(question, pool, k=2) returns the k examples from pool that best fit the question, best first. Each example is a dict: {"ticket": "...", "label": "SHIP-1"}.

How you decide is up to you. Four tools are in scope:

  • embedder.embed(texts): a real embedding model. Give it a list of strings and it returns one vector for each, in order; cosine(a, b) compares two. Texts that mean the same thing land close together even with no word in common.
  • jev.ask(state, questions): Jev, a decision model. It does not write text: you give it some content and typed questions, and it returns numbers. A score question rates the content on a rubric you write (["unrelated", "related", "same kind of problem"]), and one call can carry a question for every example. jev.choose and jev.yes_no ask a single question.
  • llm.ask(prompt): the real chat model. It can read the question and the pool and tell you which examples fit.
  • Plain Python, if you think counting shared words is enough. Customers write nothing has turned up yet; the example says my package is still not here. Try it and see.

The tests check what you picked, not how: over two sets of six tickets written the way customers write, your best example must carry the right code for at least five of each six. They also count calls: across all the tools, at most two calls per question. Embedding the question and every example in one list is one call. Embedding them one at a time is thirteen, and will fail.

few_shot_messages(question, examples) lays the prompt out. For each example, a "user" message with the ticket and then an "assistant" message with its label; last, a "user" message with the question. One detail: examples arrives best first, and the best example goes nearest the question, so lay them out in reverse. Models weigh what they read last most heavily.

Everything costs credits here: an embedding call is the cheapest, a Jev call costs more, a chat call the most, and none is free.