Skip to content

fix: read every inline part of an email body split by an attachment - #4500

Open
L4XB wants to merge 3 commits into
Unstructured-IO:mainfrom
L4XB:fix/email-split-html-body
Open

L4XB wants to merge 3 commits into
Unstructured-IO:mainfrom
L4XB:fix/email-split-html-body

Conversation

@L4XB

@L4XB L4XB commented Sep 25, 2026 •

Copy link
Copy Markdown

Summary

partition_email() reads the message body from the one part EmailMessage.get_body() returns. When the text of a message is spread over several parts of a multipart/mixed, everything after the first of them never reaches the elements. With the default content_source="text/html", that includes a common case: Apple Mail sends a message with a file placed inside its text as

multipart/alternative
├── text/plain              (the whole text, with a "<report.pdf>" placeholder)
└── multipart/mixed
    ├── text/html           ("Here is the report.")
    ├── application/pdf     (inline)
    └── text/html           ("Let me know what you think.")

get_body(preferencelist=("html", "plain")) returns the first text/html part, so the text after the file is lost.

Four layouts through partition_email() on main, with an inline file between the parts where there is one:

layout main returns this PR returns
the one above Here is the report. both sentences
mixed[html, file, html] the text before the file the text before and after
mixed[related[html, image], file, html] the text before the file the text before and after
mixed[alternative[plain, html], plain] (a mailing-list footer part) the message the message and the footer

A fifth, mixed[alternative, alternative], raises KeyError: 'multipart/alternative' on main. iter_attachments() hands over the second alternative as an attachment, and get_content() has no handler for multipart/*.

Change

unstructured/partition/email.py:

  • EmailPartitioningContext.body_parts, next to .body_part: the parts that together carry the message text, in order. It is [body_part] unless body_part is shown by a multipart/mixed part, as a direct part of it or inside one through multipart/alternative and multipart/related parts only. Then each part of that multipart/mixed that iter_attachments() would take for body text (text/plain, text/html, multipart/alternative, multipart/related) contributes part.get_body(preferencelist). That skips attached files and non-text inline parts, since get_body() returns None for them.
  • _iter_email_body_elements() partitions each of those parts, as it did the single one.
  • The attachment loop skips an attachment that is, or holds, a part read as body text. iter_attachments() passes over only the first body-like part of each kind, so a second text/html part, or a second multipart/alternative part, arrives there too. That second alternative is what raised the KeyError.
  • content_source="text/plain" is unchanged for the Apple Mail layout: the plain part is not inside the multipart/mixed, and it already carries the whole text.

A nested multipart/mixed part is not one of those body kinds, so it is left to the attachment loop as before. That is also the case #4424 handles, so the two changes do not read a part twice.

CHANGELOG.md has a 0.27.9-dev0 entry, and __version__.py is bumped to match.

Testing

In test_unstructured/partition/test_email.py, with a new fixture example-docs/eml/mime-html-split-by-inline-attachment.eml (the layout above, written for this test):

  • test_partition_email_reads_an_html_body_split_by_an_inline_attachment
  • test_partition_email_reads_every_inline_part_of_a_mixed_body: HTML parts, a related part followed by an HTML part, and a footer part.
  • test_partition_email_does_not_partition_a_body_part_again_as_an_attachment: with auto.partition mocked, a second HTML part is not handed to it as an attachment, while the inline CSV still is. With two alternative parts, nothing is handed over and nothing raises.
  • Four unit tests for .body_parts: a body in one part, the split fixture, content_source="text/plain", and a message with no body.

Results:

  • test_email.py: 82 passed. Against main's email.py, the 10 new tests fail: the 6 behaviour tests on their assertions, or with the KeyError in the two-alternatives case, and the 4 unit tests because .body_parts is new.
  • Every .eml in example-docs/eml, 40 files, partitioned with both content sources: all 80 outputs are identical to main's except for the new fixture.
  • Nine single mutations of the change, each applied alone. Eight fail at least one test:
    • reading body_part alone;
    • not looking through alternative parts, or through related parts;
    • not skipping a body part in the attachment loop, or skipping only the part itself and not a part that holds one;
    • only checking the root for multipart/mixed;
    • treating any container as multipart/mixed;
    • partitioning only the first body part.
  • The ninth drops the body-kind check. It survives, because the only difference is a nested multipart/mixed part, which the tests do not build: on main it reaches get_content() in the attachment loop and raises the KeyError fix: partition_email raises KeyError on a multipart/* attachment #4424 fixes.
  • ruff check and ruff format --check (0.15.10, the locked version) are clean on both files. CHANGELOG.md and __version__.py agree on 0.27.9-dev0. I could not run make check-version itself, because it needs GNU sed ≥ 4.3.
  • test_auto.py needs the pdf extra to import, which is not installed here. Its email cases use fake-email.eml and fake-email-attachment.eml, and both are part of the 80 identical outputs above.

To try it:

from unstructured.partition.email import partition_email

elements = partition_email("example-docs/eml/mime-html-split-by-inline-attachment.eml")
print([e.text for e in elements])
# main:     ['Here is the report.']
# this PR:  ['Here is the report.', 'Let me know what you think.']

Review in cubic

partition_email() read the body from the single part that
EmailMessage.get_body() returns. Apple Mail sends a message with a file
placed inside its text as a multipart/mixed part holding an HTML part,
the file, and another HTML part, so with the default
content_source="text/html" everything after the file was lost. A
mailing-list footer sent as a part of its own was lost the same way.

When the body part is shown by a multipart/mixed part, each inline
text, alternative or related part of it now contributes its body, in
order. iter_attachments() passes over only the first body-like part of
each kind, so an attachment that is, or holds, a part read as body text
is skipped; a body with a second multipart/alternative part therefore
no longer raises KeyError.
…body

# Conflicts:
#	CHANGELOG.md
#	unstructured/__version__.py
…body

main released 0.27.18; this entry moves to 0.27.19-dev0.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant