Gate the Tool That Sends Data Out
The agent can read private data (read_file on the customer list), read content nobody vetted (fetch_url on a web page), and send data out (http_post, send_email). Together that is the lethal trifecta: the page tells the model to post the customer list somewhere, and it does. The fix is to make sending unavailable once untrusted content has arrived.
Implement run_gated(task, model, tools, untrusted, sinks, max_turns=8): the standard tool loop with a taint rule. untrusted names the tools whose results are untrusted content; sinks names the tools that send data out.
- Once a tool in
untrustedhas been called, the context is tainted for the rest of the run. Nothing later untaints it. - While tainted, a call to a tool in
sinksis refused: it does not run. The model gets anis_errortool result whose content starts withblocked:. Record the tool's name inblocked. - Before any untrusted content has arrived, sinks run normally. Sending a report and then reading a web page is fine.
- Every other tool always runs. Reading private data is not the problem; sending it is.
- Return
{"answer", "turns", "blocked", "tainted"}.answeris the model's final text, orNoneifmax_turnsran out.
The scripted model does what the poisoned page says: fetches the page, reads customers.csv, and tries to post it. With the gate the post is refused and the fixture's outbox stays empty.