Try the Small Model First
The large model answers everything well and costs ten times more; the small model answers most things well. Ask the small one first and escalate only what it gets wrong.
Implement cascade(question, small, large, accept) returning {"answer": str, "escalated": bool}.
- Ask
small. Ifaccept(answer)returnsTrue, you are done; never touchlarge. - Otherwise ask
largeand return its answer withescalated: True.
accept is the gate, and it must be cheap and deterministic (a format check, a validator, a regex): a model deciding whether the small model was right saves no call. In these tests the gate checks the answer is a bare number.
The tests script small so each path is reproducible (one script answers acceptably, one with prose the gate rejects); large is real. You are scored on cost, so both models' call counts are checked.