Skip to content

< all problems41 · Level 07, Evals

Catch the Regression Between Two Versions

medium · implement · Evals

Your eval pass rate went from 84% to 86% after a change. That is compatible with specific cases going from pass to fail: better at the easy things, broken on something a real user depends on, and the average hid it.

Implement compare_runs(baseline, candidate) where each run is a dict of case_id -> {"passed": bool, "score": float}. Return:

{"regressed": [...], "improved": [...], "missing": [...],
 "pass_rate_delta": float, "verdict": "regression" | "ok"}
  1. regressed: cases that passed in baseline and fail in candidate. Sorted, so the report is stable.
  2. improved: the reverse.
  3. missing: cases in the baseline that the candidate did not run. A case that vanished is not a case that passed.
  4. pass_rate_delta: candidate pass rate minus baseline pass rate, over the cases both ran, rounded to 4 places.
  5. verdict: "regression" if any case regressed or any case is missing, regardless of the aggregate. Otherwise "ok".

A single per-case regression fails the verdict even when the average improved.