Repository navigation
fix(sanitizer): redact a sensitive KEY=value chained after a harmless pair (#3311) - #3335
Conversation
… pair (#3311) #1670's single KEY=value regex took the key up to the first `=` and the value up to whitespace, so in `user=a&password=X` the harmless `user` pair consumed `a&password=X` and the sensitive pair was never examined (same for `?client=me&access_token=X` and `--env=GITHUB_TOKEN=X`). The pre-#1661 regex redacted all of these. _KV_LINE_RE now matches only `KEY=` (the key also stops at `&`, `;` and `,`), and _redact_kv_pairs walks the keys in order. A harmless key consumes nothing, so the next key in the chain is still checked. Only a sensitive key takes its value, in the old up-to-whitespace shape, so a secret containing a separator is never split and partly leaked, and the Google consent-link exemption still sees the whole URL. The search only moves forward and the lookbehind is kept, so the pass stays linear. Applied identically to the agent-server copy. Red with the fix reverted: TestChainedPairs (both copies) and test_ec_input_hardening_edges::TestSanitizerBoundaries:: test_a_sensitive_pair_chained_after_a_harmless_one_is_redacted (xfail removed). Fixes #3311 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| pos = key.end() | ||
| if not _is_sensitive_kv_key(key.group(1)): | ||
| continue | ||
| pair = _KV_PAIR_RE.match(text, key.start()) |
|
merge-train: ejected from this batch — rides the next train once fixed. Over-redaction vs #3311 AC1 ("harmless pairs are kept"). The CodeQL #383/#384 (py/polynomial-redos, Minor: a differential fuzz found |
… sensitive word (#3311) Merge-train finding on #3335: walking the KEY=value chain put keys under the containment key rule (`.*AUTH.*`, `.*TOKEN.*`, ...) that dev never examined, because they sat inside a harmless pair's value. Everyday query parameters were redacted, and since a sensitive value runs to whitespace each one took the rest of the URL with it: .../issues?q=is:open&author=bob&page=2 ?page=2&tokens_used=10&model=x ?sort=asc&passwordless=true&lang=en AC1 says the harmless pairs are kept. For a chained key only, a closed list of harmless words is blanked out before the same containment test: AUTHOR... (not authoriz/authoris), PASSWORDLESS, SECRETARY/-IES/-IAT, TOKENIZ/TOKENIS, and the token-count names (max/total/input/output/prompt/completion_tokens, token(s)_used/_count/_limit/_usage). A word list rather than a word-boundary rule, because a boundary rule leaks an open class (secretkey, authkey, tokenvalue, passwords, credentials, authorization); with the list, the remainder is still tested, so author_token and max_tokens_secret still redact. "Chained" means the key sits inside what #1670 consumed as a harmless pair's value: after `=`, `&`, `;` or `,`, with only value characters back to the previous `KEY=`. Every key dev already tested - one starting its run, following a bare word, or following a redacted pair - keeps plain containment, so this never redacts less than dev did. The gap check scans only the text between two consecutive keys, so the walk stays linear. Both copies (backend and agent-server) carry the identical change. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Addressed the merge-train over-redaction finding in Rule. For a chained key only (one sitting inside what Now kept, as on Known limits, stated plainly.
Evidence. Red first: 26 failed on the previous head with the new cases. After: |
| pos = 0 # where the next key search starts; only ever moves forward | ||
| prev = -1 # end of the previous harmless `KEY=`; -1 = none in this chain | ||
| while True: | ||
| key = _KV_LINE_RE.search(text, pos) |
vybe
left a comment
There was a problem hiding this comment.
merge-train: validated at lane C (/validate-pr, /review, /cso --diff). CodeQL alerts 388/384 measured linear on 19 adversarial shapes up to 400k chars; left open for a human dismissal.
|
merge-train: merged. One follow-up noted from validation, accepted at the merge gate: when |
Summary
Since #1670,
sanitize_text's KEY=value pass took the key up to the first=and the value up to the next whitespace. Inuser=a&password=Xthe harmlessuserpair swalloweda&password=X, so the sensitive pair was never checked and the string came back unredacted. The same happened with?client=me&access_token=Xand--env=GITHUB_TOKEN=X.Fix (identical in the backend and agent-server copies):
_KV_LINE_REnow matches onlyKEY=; a key also stops at&,;and,._redact_kv_pairswalks the keys in order. A harmless key consumes nothing, so the next key in the chain is still checked._KV_PAIR_RE), then goes through the unchanged_redact_kv_match.Behaviour to review: a sensitive value still runs to the next whitespace, so
password=x&b=2becomespassword=***REDACTED***and also hides&b=2. This is deliberate:&,;or,is never split and partly leaked;Harmless pairs before the secret are kept as they were.
test_a_sensitive_value_is_never_split_at_a_separatorpins this.Tests
test_ec_input_hardening_edges.py::…::test_a_sensitive_pair_chained_after_a_harmless_one_is_redacted: strict-xfail marker removed.TestChainedPairsintest_1661_sanitizer_linear.py, run against both copies. It covers:&,;,,, space, tab;test_1661,test_2398,test_google_consent_sanitizer): 205 passed.Fixes #3311
🤖 Generated with Claude Code