Skip to content

Scan XML element text and attributes at any depth - #730

Open
iacobdaniel wants to merge 3 commits into
mainfrom
fix/scan-xml-element-text
Open

iacobdaniel wants to merge 3 commits into
mainfrom
fix/scan-xml-element-text

Conversation

@iacobdaniel

Copy link
Copy Markdown
Contributor

Only the attributes of the XML elements directly under the root were read. The whole document is now walked, storing element text as well as attributes, at any depth.

Comment thread aikido_zen/helpers/extract_data_from_xml_body.py Outdated
Comment on lines +25 to +30
for text in (element.text, element.tail):
if text:
stripped = text.strip()
if stripped:
extracted_xml.setdefault(element.tag, set()).update(
(text, stripped)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Medium - Ignored XML text is treated as the outbound URL source

A caller can send XML containing a private-service hostname in text or tail content that the application never uses, while the handler makes a fixed or independently derived request to that service. The new traversal adds that ignored fragment to context.xml, and the SSRF checker scans every XML value as though it tainted the destination. With blocking enabled, the request is rejected as SSRF and the legitimate outbound operation does not run.

Show fix

Only associate XML values that flow into the outbound hostname (or preserve a structured XML path for sink-aware matching); do not promote every text/tail fragment into a global source set for SSRF matching.

More info - Reply on this comment to give feedback or ignore the issue.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At the outbound call the SSRF check only sees that the destination is private and that the hostname appears somewhere in the request — it can't tell whether the app took that value or picked the destination itself, so both get reported. That already happens on main when the URL sits in an XML attribute, a JSON field or a query param; this PR only adds element text as one more place such a string can sit.

@iacobdaniel

Copy link
Copy Markdown
Contributor Author

Not ready to merge yet. Scanning the text as well as the attributes roughly doubles what is stored per request, and nothing caps that. Is that cost acceptable, or should the amount extracted be limited by size?

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant