The environment is part of the problem
One problem gives you stdlib and nothing else. The next gives you PyTorch and a real Hugging Face model. A third gives you a live model that returns malformed JSON, because that is what live models do.
Every problem is one you implement yourself: the loop, the tools, the retrieval, the evals. The model is scripted, so your code is the only variable. Real Python, no API key, free.
Free. No account to try.
calls0/3
tokens0/1200
One agent run against a scripted model. Break a tool, drop a result, cut the budget: an agent is judged by what it does when things go wrong.
Deterministic by designThe model replays a script, like a test fixture. No flaky grading, no spend, and a failure means your code, not the dice.
Not only implementDebug a loop that works until a tool raises. Optimize one that is correct but over its call, token and dollar budget.
Sixteen tracks from a first API call to agents that survive production. Each one is a set of concepts you implement, debug and optimize, not videos you watch. The first 4 are open now; the rest are coming soon.
One problem gives you stdlib and nothing else. The next gives you PyTorch and a real Hugging Face model. A third gives you a live model that returns malformed JSON, because that is what live models do.
A RAG pipeline that runs, raises nothing, and retrieves the wrong documents. Three real defects. Find them. That is closer to the job than any implementation exercise.
Everyone who solves it passes the tests, so passing separates nobody. Where it means something you are ranked on model calls, tokens, documents read: the things that show up on a bill.
All three run in the page, against a scripted model. Parse a reply is problem 1 in the catalogue; the others are samples of the implement and debug work in the curriculum.
Every agent is a loop around one condition: keep going while the model is asking for a tool. Fill it in and run it against three recorded runs.
def run(prompt): messages = [user(prompt)] while True: resp = llm(messages) if resp.stop_reason != "": return resp.text messages += [assistant(resp), *call_tools(resp)]
Fill the blank, then Run.
The first problem takes five minutes and runs without an account.
Try the first problem