Score an Answer with an LLM Judge
Have a model grade an answer against a rubric.
Implement judge(llm, question, answer) returning an integer from 1 to 5.
- Put the rubric in the prompt, 1 for useless up to 5 for excellent. A judge told only "score this answer" invents a new scale on every call.
- Find the digit. The judge will happily reply
"I'd rate this a solid 4 out of 5 because...". Return aninthowever it phrases things. - Clamp into 1 to 5. If parsing fails or the model returns 9, an out-of-range score would skew every average downstream.
A real model answers, so the tests check properties: an int in range, a good answer outscoring a nonsense one, a correct answer scoring at least 4.