Skip to content

Implement AI agents from scratch.

Every problem is one you implement yourself: the loop, the tools, the retrieval, the evals. The model is scripted, so your code is the only variable. Real Python, no API key, free.

Free. No account to try.

agent.run("what is hnsw?")
running

    calls0/3

    tokens0/1200

    One agent run against a scripted model. Break a tool, drop a result, cut the budget: an agent is judged by what it does when things go wrong.

    Deep-Agnt, from the team behind Deep-MLWrite the functionDrive the agent loopRepair a repo by prompting a coding agentImplement · Debug · Optimize · Design · Repair · Investigate · Agent · Write testsGraded against a real modelLLM APIs → Search → RAG → Tool Calling → Agents → MCP → Evals → ProductionNo API keys needed

    From the team behind Deep-ML.

    100k+engineers on Deep-ML, one account
    85problems live
    16curriculum tracks
    5challenge types
    0API keys needed

    Deterministic by designThe model replays a script, like a test fixture. No flaky grading, no spend, and a failure means your code, not the dice.

    Not only implementDebug a loop that works until a tool raises. Optimize one that is correct but over its call, token and dollar budget.

    What you learn, in order.

    Sixteen tracks from a first API call to agents that survive production. Each one is a set of concepts you implement, debug and optimize, not videos you watch. The first 4 are open now; the rest are coming soon.

    1. 01Foundations4 tracks
    2. 02Agents5 tracks · soon
    3. 03Connections2 tracks · soon
    4. 04Engineering4 tracks · soon
    5. 05Production1 tracks · soon
    Walk the curriculum →

    Why Deep-Agnt.

    The environment is part of the problem

    One problem gives you stdlib and nothing else. The next gives you PyTorch and a real Hugging Face model. A third gives you a live model that returns malformed JSON, because that is what live models do.

    Debug, not just implement

    A RAG pipeline that runs, raises nothing, and retrieves the wrong documents. Three real defects. Find them. That is closer to the job than any implementation exercise.

    Compared on what matters

    Everyone who solves it passes the tests, so passing separates nobody. Where it means something you are ranked on model calls, tokens, documents read: the things that show up on a bill.

    What you actually implement.

    All three run in the page, against a scripted model. Parse a reply is problem 1 in the catalogue; the others are samples of the implement and debug work in the curriculum.

    Close the loop.

    Every agent is a loop around one condition: keep going while the model is asking for a tool. Fill it in and run it against three recorded runs.

    loop.pyunsolved
    def run(prompt):    messages = [user(prompt)]    while True:        resp = llm(messages)        if resp.stop_reason != "":            return resp.text        messages += [assistant(resp), *call_tools(resp)]
    • run("what is hnsw?") → 'HNSW is a hierarchical graph index.'
    • run("capital of France?") → 'Paris.'
    • run("go") → 'done'

    Fill the blank, then Run.

    Start building.

    The first problem takes five minutes and runs without an account.

    Try the first problem