Skip to content

< all problems23 · Level 08, Production

Stop Paying for the Same Answer Twice

medium · implement · Production

The same question gets asked over and over, and every repeat is a model call you paid for when you already had the answer.

Implement CachedLLM, wrapping a model:

  1. ask(prompt) returns the model's answer.
  2. The same prompt asked twice makes one model call, not two.
  3. hits counts how many times the cache answered.
  4. Normalise the key. "What is RAG?" and "what is rag? " are the same question. Case and whitespace should not cost a call, but normalise too aggressively and you return the wrong answer to a question that only looked similar.
  5. Bound the size. Keep at most max_entries, evicting the oldest first.

You are graded on correctness and on model_calls. Two solutions can both return the right answers while one costs twice as much.