Problems.
Write the piece, drive the whole agent loop, or repair a repo by prompting a coding agent. Most are graded against a real model, because that is what you will ship against.
0 of 85 solved
Show me at , . 85 matches
- 16Score an Answer with an LLM Judgeimplement · liveagent · liveL07medium
- 17Ignore the Instructions Hidden in Your Documentsdebug · liveagent · liveL08hard
- 18Answer an MCP HandshakeimplementstdlibL06easy
- 19Serve a Tool Call over MCPimplementstdlibL06medium
- 20Never Reply to an MCP NotificationdebugstdlibL06medium
- 21Dispatch a Model's Tool CallimplementagentL04easy
- 22Refuse the Arguments a Model InventedimplementagentL04medium
- 23Stop Paying for the Same Answer Twiceimplement · liveagent · liveL08medium
- 24Measure Retrieval with Recall@k and MRRimplementnumpyL02medium
- 25Route a Question to the Right Toolimplement · liveagent · liveL04medium
- 26Make the Model Admit It Doesn't Knowimplement · liveagent · liveL03hard
- 27Score a Model Against a Golden Setimplement · liveagent · liveL07medium
- 28Answer Over a Context That Does Not Fitimplement · liveagent · liveL08hard
- 29Build a LangGraph With a Conditional Edgeimplement · livelanggraphL05medium
- 30Fix a LangGraph That Forgets Everythingdebug · livelanggraphL05hard
85 problems, showing 16 to 30