Catch the Regression Between Two Versions
Your eval pass rate went from 84% to 86% after a change. That is compatible with specific cases going from pass to fail: better at the easy things, broken on something a real user depends on, and the average hid it.
Implement compare_runs(baseline, candidate) where each run is a dict of case_id -> {"passed": bool, "score": float}. Return:
{"regressed": [...], "improved": [...], "missing": [...],
"pass_rate_delta": float, "verdict": "regression" | "ok"}
regressed: cases that passed in baseline and fail in candidate. Sorted, so the report is stable.improved: the reverse.missing: cases in the baseline that the candidate did not run. A case that vanished is not a case that passed.pass_rate_delta: candidate pass rate minus baseline pass rate, over the cases both ran, rounded to 4 places.verdict:"regression"if any case regressed or any case is missing, regardless of the aggregate. Otherwise"ok".
A single per-case regression fails the verdict even when the average improved.