Stop Paying for the Same Answer Twice
The same question gets asked over and over, and every repeat is a model call you paid for when you already had the answer.
Implement CachedLLM, wrapping a model:
ask(prompt)returns the model's answer.- The same prompt asked twice makes one model call, not two.
hitscounts how many times the cache answered.- Normalise the key.
"What is RAG?"and"what is rag? "are the same question. Case and whitespace should not cost a call, but normalise too aggressively and you return the wrong answer to a question that only looked similar. - Bound the size. Keep at most
max_entries, evicting the oldest first.
You are graded on correctness and on model_calls. Two solutions can both return the right answers while one costs twice as much.