Skip to content

< CurriculumModel APIs · 14 of 46 ·38 · Level 01, LLM APIs

Batch Calls Without Tripping the Rate Limit

medium · implement · Production

Send fifty requests to a model provider in a tight loop and around the tenth you get a 429. Implement two functions: one retries a single call with capped, jittered backoff, the other runs a batch through it.

ask_with_backoff(llm, prompt, max_retries=5, base_delay=1.0, max_delay=30.0, jitter=None, sleep=time.sleep)

  1. On a RateLimitError or OverloadedError, wait and try again. Wait exponentially longer each time: base_delay, then double it, then double again.
  2. Cap the wait at max_delay. No single wait may be longer than the cap, however many retries came before it.
  3. Jitter. jitter, when given, is a function returning a number in [0, 1). Call it once per wait and wait delay * (0.5 + 0.5 * that number), where delay is the capped delay, so each wait lands between half the delay and all of it. With jitter=None, wait the delay itself.
  4. Any other error is your request being wrong, not the provider being busy. Re-raise it immediately.
  5. After max_retries retries, give up and raise the last error.
  6. sleep and jitter are injectable. Tests pass a recorder and a fixed sequence to check what you waited without waiting.

ask_all(llm, prompts, **kwargs) runs every prompt through the first function and returns the replies in order. Two kinds of failure, two different answers:

  • A prompt whose request is wrong (any APIError that is not one of the two busy signals) fails alone: put None in its place and carry on.
  • A prompt that runs out of retries means the provider is still down: let that error out and stop.

The model is scripted to fail on cue. Your loop is the only thing under test.