Skip to content

< all projects28 steps · hard · Rebuild

Multi-harness RL

You will have built the training side of agentic RL as Hugging Face's guide describes it: the same model scores very differently depending on the harness around it, models trained inside one harness overfit to its conventions, so train inside several at once, with a capture proxy between any harness and the model. Six parts, 28 functions, standard library only. Tokens, not text: why a trainer needs ids and log probabilities and what re-tokenization does to a sample. The capture proxy: tell the request dialects apart, convert between Messages and chat completions both ways, grade an engine's capture level, record a call. The rollout graph: link calls by longest exact token prefix, cut root-to-leaf paths into sequences, mask what was not sampled, refuse to train on a session with no tokens, and cross-check against the harness's own trace. Tasks and rewards: a task that names no harness, one number per rollout, the efficiency bonus, and the reward hack that fooled a curve. GRPO: group advantages, importance ratios and the clipped objective on numbers you can check. Across harnesses: the pass matrix, saved calls, scaffold reversal and transfer loss. The ending runs your functions twice. First on a scripted full-capture engine, where eight rollouts become a real batch with advantages and a clipped objective. Then around the site's live model, where your proxy records every call and your batch function refuses, because a hosted API returns no token ids. That refusal is the guide's point, and you wrote it. What is not here: a GPU, model weights, or Claude Code in a container. The optimizer step and the real harnesses are cited, not rebuilt.

Start: Tokenize by longest matchSign in to run steps and keep your progress.
  1. Part 01 · 5 steps

    Tokens, not text

    What a trainer needs from a model call, and why decoded text is not it.

    1. 01Tokenize by longest matchnext
    2. 02Decode ids
    3. 03Catch re-tokenization drift
    4. 04The record a trainer can use
    5. 05Refuse truncated sampling
  2. Part 02 · 5 steps

    The capture proxy

    Sit between any harness and the model; speak its dialect; record what training needs.

    1. 06Tell the dialects apart
    2. 07Convert a Messages request to chat
    3. 08Replay the answer in the caller's format
    4. 09Grade an engine's capture level
    5. 10Record one call
  3. Part 03 · 6 steps

    The rollout graph

    Link calls by their token prefix, cut paths into sequences, mask what was not sampled.

    1. 11Find a call's parent
    2. 12Build the graph
    3. 13Every root-to-leaf path
    4. 14Cut a path into a sequence
    5. 15The batch, or a refusal
    6. 16Cross-check against the harness's own trace
  4. Part 04 · 5 steps

    Tasks and rewards

    A task that names no harness, one number per rollout, and the hack that fools a curve.

    1. 17A task that names no harness
    2. 18One number per rollout
    3. 19The efficiency bonus
    4. 20Reward a rollout
    5. 21Catch the reward hack
  5. Part 05 · 3 steps

    GRPO without a GPU

    Group advantages, importance ratios and the clipped objective, on numbers you can check.

    1. 22Group advantages
    2. 23Importance ratios
    3. 24The clipped objective
  6. Part 06 · 4 steps

    Across harnesses

    Read a results matrix: pass rates, saved calls, and when the best model depends on the harness.

    1. 25The pass matrix
    2. 26Saved tool calls
    3. 27Scaffold reversal
    4. 28Transfer loss