Skip to content

< all problems03 · Level 01, LLM APIs

Retry a Failed Model Call

easy · implement · LLM Fundamentals

Model APIs fail. They rate-limit you, they time out, they return a 503 because a GPU somewhere fell over. Almost all of it is transient, and the difference between a flaky product and a reliable one is a retry loop that knows what's worth retrying.

Implement call_with_retry(fn, max_attempts=3, base_delay=1.0).

Call fn(). If it succeeds, return its result. If it raises:

  1. Retry RateLimitError, TimeoutError, and ServerError — these are transient.
  2. Do not retry anything else. A ValueError from a malformed request will fail identically forever; retrying just wastes time and money.
  3. Wait base_delay * (2 ** attempt) before each retry — 1s, 2s, 4s — using the provided sleep(). Hammering a rate-limited endpoint makes it worse.
  4. After max_attempts total attempts, re-raise the last exception.

Use the provided sleep() rather than time.sleep so the tests don't actually wait.