1 file · saved in this browser
Your agent.
The coding agent that drives every repair problem. It starts as the simplest one that works; make it yours, any shape you like, as long as harness.py defines run_coding_agent.
You are talking to the agent in the editor, run for real.
It can work on a small practice repo, read its own docs, and edit its own code when you ask. You decide whether to keep those edits. It never sees a problem's repo here; it meets those in the graded run on the problem page.
Scripted models, a fixture repo, no key and no credits. The check is a model that behaves: pass it and repair problems can run on this agent. The bench is five that do not, 8 calls each: it shows what to improve. Your tests are whatever you want to hold it to.
Your tests
None yet. Write tests of your own: functions named test_ in a file named test_*.py, run on a model you script, so they cost nothing. They run with every Test, next to the bench.
History
Every version you tested, with how it did. Saved in this browser, the last twelve.
Nothing yet. Test your agent and that version is kept here.