Skip to content

fix(html): extract <figure> content instead of discarding it - #4503

Open
r0h1tb wants to merge 2 commits into
Unstructured-IO:mainfrom
r0h1tb:fix/html-figure
Open

r0h1tb wants to merge 2 commits into
Unstructured-IO:mainfrom
r0h1tb:fix/html-figure

Conversation

@r0h1tb

@r0h1tb r0h1tb commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

Problem

<figure> is mapped to RemovedBlock, so partition_html() drops every figure together with its image and <figcaption>. Wikipedia wraps every thumbnail in one, which is the case #3606 reports:

partition_html(text=(
    '<p>Intro.</p><figure><a href="/wiki/File:Tonnetz.svg"><img src="//upload.wikimedia.org/tonnetz.svg" alt=""></a>'
    "<figcaption>The Tonnetz, minor as upside-down major</figcaption></figure>"
))
# [Text('Intro.')]

Fix

A figure is now traversed like a <div>: its <img> becomes an Image element with image_url, and a code listing or table inside it is kept. <figcaption> gets its own block class that emits a FigureCaption, including when the caption is wrapped in a single <p>.

On #3606 @scanny suggested putting the caption into Image.text. I went with a separate FigureCaption element instead, because it keeps the image's alt text, works for figures that aren't images, and matches what hi_res PDF partitioning emits for a picture and its caption. Happy to switch if you'd rather have it on the image.

Figure is its own class rather than plain Flow so that ListItemBlock doesn't adopt a text-only figure as a loose-list paragraph.

An Image or CodeSnippet from a figure used to pop the active heading in set_element_hierarchy(), since neither was a valid child of Title or Header, so the caption and everything after it lost their section parent_id. Both are now allowed under Title and Header. That list is shared, so this also applies to images and code from other partitioners, which had the same problem. I checked the stored ingest outputs and none of them change.

Flow._page_number stopped at the first non-Flow parent, so an image inside an <a> had no page number. It now checks every ancestor. A caption adopted from a sole <p> takes that <p>'s page, as ListItemBlock already does.

Tests

New tests cover the MediaWiki markup above (element types, order, image_url and the caption's link annotation), a code listing and a table inside a figure, and a <p>-wrapped caption. All four fail on main with the figure missing entirely. The html, chunking, documents, md, text and email tests go from 1201 passed on main to 1205, with no failures.

test_partition_html_keeps_the_heading_as_parent_across_a_figure (image and code variants) checks parent_id for the figure content, the caption, the paragraph after it and a following subsection. test_partition_html_gives_figure_elements_their_own_page_number covers an <a>-wrapped image and a caption whose page overrides its container. All three fail without the second commit. After rebasing: html, common, documents, chunking and staging go from 1154 passed to 1157, 41 skipped; four modules need extras that aren't installed here, same on main. Ruff is clean.

Two parser tests used <figure> as their example of a removed block; they now use <nav>, which is still removed.

Rebased on main, version 0.27.19. #4451 bumps to the same version, so whichever lands second needs it moved.

Resolves #3606.

Review in cubic

@cragwolfe cragwolfe left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

High — Figure content drops the active heading from hierarchy

Location: unstructured/partition/html/parser.py:1069-1071 (figure/figcaption registration), with traversal in :462-489.

The new traversal emits Image and CodeSnippet elements from figures, but the hierarchy rules allow neither category beneath Title or Header (unstructured/partition/common/metadata.py:58-82). set_element_hierarchy() pops incompatible stack entries (:142-167), so an image or code listing removes the active heading. For <h1>Section</h1><figure><img ...><figcaption>Diagram</figcaption></figure><p>After</p>, the image clears the heading; the caption and following paragraph then receive no section parent_id. The same occurs for a figure containing <pre>. This silently breaks heading-based section reconstruction for documents with figures.

Preserve the active heading across non-structural figure leaves (for example, make Image and CodeSnippet valid heading children or adjust stack handling). Add parent_id assertions for image and code figures, captions, following paragraphs, and a subsequent subsection.

High — Figure-derived elements can have missing or incorrect page numbers

Location: unstructured/partition/html/parser.py:475-489.

For the motivating anchor-wrapped image shape, ImageBlock reads its own _page_number (:552-577), but Flow._page_number only follows a parent when it is another Flow (:352-364). An <a> is phrasing, so an image inside <div data-page-number="7"><figure><a><img ...></a>... gets page_number=None, while the caption inherits page 7. The new single-<p> caption adoption path also builds the FigureCaption from the <figcaption> accumulator without applying the child block's page number; a child <p data-page-number="8"> is therefore assigned the outer page 7. ListItemBlock handles this same adoption case by copying child._page_number (:527-531).

Walk all ancestors when resolving page metadata, and copy the adopted child block's page number onto the emitted caption. Add focused assertions for an anchor-wrapped image and a wrapped caption whose page overrides its container. Incorrect page metadata also changes deterministic element IDs because page number and per-page sequence are hash inputs (unstructured/partition/common/metadata.py:306-327).

(authored by codex)

r0h1tb added 2 commits October 5, 2026 20:35
<figure> was mapped to RemovedBlock, so partition_html() dropped every
figure along with its image and caption, including every Wikipedia
thumbnail.

Traverse a figure like a <div>, so its <img> becomes an Image element
and a code listing or table inside it is kept, and give <figcaption>
its own block class that emits a FigureCaption, including when the
caption is wrapped in a single <p>. Figure is a distinct class so a list
item does not adopt one as a plain paragraph.

Two parser tests used <figure> as their example of a removed block and
now use <nav> instead.
An Image or CodeSnippet from a figure was not a valid child of Title or
Header in HIERARCHY_RULE_SET, so set_element_hierarchy() popped the
active heading: the caption and everything after it lost their section
parent_id. Allow both under Title and Header.

Flow._page_number stopped at the first non-Flow parent, so an image
inside an <a> got no page number. Walk every ancestor instead. A caption
adopted from a sole <p> in <figcaption> now takes that <p>'s page, as
ListItemBlock already does.
@r0h1tb

r0h1tb commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Both fixed in 5036f11.

  • Hierarchy: Image and CodeSnippet are now valid children of Title and Header in HIERARCHY_RULE_SET, so an image or code listing no longer pops the heading. That list is shared, so it applies to the other partitioners too. I checked the stored ingest outputs and none of them change. The tests check parent_id for image and code figures, the caption, the paragraph after, and a following subsection.
  • Page numbers: Flow._page_number now walks every ancestor, so the <a>-wrapped image gets page 7, and the adopted <p> caption keeps its own page 8, same as ListItemBlock.

Also rebased on main, now 0.27.19. #4451 bumps to the same version, so whichever goes in second needs it moved.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat: capture images/figures

2 participants