Skip to content

< all problems56 · Level 08, Production

Redact Secrets Before They Reach the Model

medium · implement · Production

Ticket from platform: our API key is in a trace. The agent ran cat .env, the file went into the context, and the context went to the trace exporter. Filter the tool result before the model sees it.

Implement redact(text) returning (clean_text, counts). Replace each secret with [REDACTED:<kind>] and count the replacements per kind in a dict. Six kinds, applied in this order, each on the output of the one before:

  1. private_key: a PEM block, from -----BEGIN ... PRIVATE KEY----- to the matching -----END ... PRIVATE KEY-----, lines between included. One placeholder for the whole block.
  2. url_credentials: the user:password between :// and @ in a URL. Keep the scheme and the host.
  3. api_key: sk- followed by 20 or more of [A-Za-z0-9_-]; AKIA followed by 16 uppercase letters or digits; ghp_ followed by 36 letters or digits.
  4. bearer: Bearer followed by 20 or more of [A-Za-z0-9._-]. Keep the word Bearer.
  5. env_secret: NAME=value where NAME is uppercase letters, digits and underscores ending in KEY, SECRET, TOKEN or PASSWORD. Keep the name; replace the value, which runs to the next whitespace.
  6. password: password or passwd in any case, then optional spaces, : or =, optional spaces, and a value up to the next whitespace. Keep the label.

Return counts only for the kinds that occurred.

The catch: a value that is already a placeholder is left alone, so OPENAI_API_KEY=sk-... becomes OPENAI_API_KEY=[REDACTED:api_key] under rule 3 and rule 5 does not touch it again; redact applied twice is redact applied once. And sk-test is not a key: the length thresholds keep ordinary text intact.