Skip to content

fix(json): preserve rehydrated element identity - #4504

Open
rohan-patnaik wants to merge 4 commits into
Unstructured-IO:mainfrom
rohan-patnaik:rohan-patnaik/preserve-json-metadata
Open

rohan-patnaik wants to merge 4 commits into
Unstructured-IO:mainfrom
rohan-patnaik:rohan-patnaik/preserve-json-metadata

Conversation

@rohan-patnaik

@rohan-patnaik rohan-patnaik commented Sep 27, 2026 •

Copy link
Copy Markdown

Summary

Preserve serialized element IDs and source metadata when partition_json() rehydrates Unstructured output. Generic JSON metadata and deterministic ID processing now apply only to arbitrary JSON, while explicit metadata overrides remain supported.

Fixes #3365.

Verification

  • Added a fail-before/pass-after full element-dictionary round trip
  • Passed all 60 applicable JSON partition tests
  • Passed whole-tree Ruff lint and formatting checks
  • Passed changelog and package-version synchronization

One existing deep-recursion expectation is excluded because it fails identically on current main under Python 3.13.

Review in cubic

@cragwolfe

Copy link
Copy Markdown
Contributor

Review findings

I found three Medium-priority compatibility regressions introduced by this refactor. The core unchunked rehydration behavior—preserving serialized element IDs and stored metadata—looks correct.

[Medium] Preserve last_modified on arbitrary-JSON orig_elements

Location: unstructured/partition/json.py:111-112,130-138

The arbitrary-JSON branch now calls _apply_chunking() before assigning last_modified. Built-in chunkers retain the pre-chunk elements in chunk.metadata.orig_elements, so the returned chunk receives the timestamp but its retained original elements do not. For example, with metadata_last_modified="2020-07-05T09:24:28", chunk.metadata.last_modified is populated while chunk.metadata.orig_elements[0].metadata.last_modified is None. This is a provenance regression for consumers that inspect or serialize orig_elements.

Suggested fix: assign metadata_last_modified or last_modified to the raw arbitrary elements before _apply_chunking(), and remove the replacement assignment from the final arbitrary-element loop. Keep this fix limited to arbitrary JSON so rehydrated elements continue to preserve their stored timestamps.

[Medium] Do not rename rehydrated attachment elements

Location: unstructured/partition/json.py:115-118

The removed metadata decorator skipped filename overrides when element.metadata.attached_to_filename was set. The new rehydration loop applies metadata_filename to every element, so a serialized attachment with filename="invoice.pdf" is renamed to the top-level override (for example, renamed.eml). That destroys the attachment source filename while leaving its attachment relationship and MIME type intact.

Suggested fix: retain the old guard for the filename override: if metadata_filename and element.metadata.attached_to_filename is None: before calling add_element_metadata(..., filename=metadata_filename). Keep explicit metadata_last_modified handling separate because it historically applied to all top-level elements.

[Medium] Preserve complete bound arguments for registered chunkers

Location: unstructured/partition/json.py:31-38

The old add_chunking_strategy decorator used get_call_args_applying_defaults() to pass the complete bound partition call to registered chunkers. _apply_chunking() now forwards only kwargs, so named parameters such as filename, file, text, and metadata_last_modified are missing. A custom registered chunker that requires filename or metadata_last_modified now raises TypeError; chunkers with defaults silently receive the wrong values. Built-in basic and by_title do not expose this because they do not consume those arguments.

Suggested fix: construct the equivalent complete call-argument dictionary—including named parameters, defaults, and kwargs—at every _apply_chunking() call site, including the early empty-input returns, then let chunk() filter it against the registered chunker signature.

(authored by codex)

@rohan-patnaik
rohan-patnaik force-pushed the rohan-patnaik/preserve-json-metadata branch from 2b5617d to 9e42507 Compare September 27, 2026 20:15
@rohan-patnaik

Copy link
Copy Markdown
Author

Addressed all three compatibility findings in 9e42507b and rebased onto current main:

  • arbitrary-JSON elements now receive last_modified before chunking, so retained orig_elements preserve it;
  • rehydrated attachment elements keep their own filename when metadata_filename is supplied;
  • chunk dispatch now receives the complete bound partition_json call, including explicit/default arguments on both empty-input paths.

I also advanced the development version to 0.27.11 to resolve the upstream 0.27.10 collision. Added regressions for each boundary. Validation: 64 JSON partition tests, repository Ruff/version checks, and targeted mypy all pass.

Signed-off-by: Rohan Patnaik <rohan-patnaik@users.noreply.github.com>
@rohan-patnaik

Copy link
Copy Markdown
Author

Resolved the new release conflicts in 09d7669: preserved main’s changelog entries and advanced this change to 0.27.13. The JSON implementation is unchanged. All 64 JSON tests, Ruff lint/format, and version-sync validation pass; GitHub now reports the PR mergeable.

@cragwolfe cragwolfe left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed head 09d7669e5d8c0b022d06796a89a8433070551206 against base 74e2fea06836169ee7749d61ce7d555ce4b83051, including the full production Oracle Pro review and independent source verification. The three earlier findings are fixed: arbitrary-JSON originals retain timestamps, attachment filenames survive an explicit filename override, and registered chunkers receive the complete bound arguments on all four dispatch paths.

[Medium] Apply explicit timestamp overrides before chunking rehydrated elements

Affected code: unstructured/partition/json.py:142-155.

The rehydration branch calls _apply_chunking() before applying metadata_last_modified. Built-in chunkers create separate chunk metadata and retain the input elements in metadata.orig_elements (unstructured/chunking/base.py:878-880,961-977; tables retain a separate copy at 1185-1188). Updating the returned chunk at line 155 therefore leaves its originals with missing or stale timestamps.

For example:

chunks = partition_json(
    text='[{"type":"NarrativeText","element_id":"source-id",'
         '"text":"Saved document content."}]',
    metadata_last_modified="2020-07-05T09:24:28",
    chunking_strategy="basic",
)

The returned chunk receives the explicit timestamp, but chunks[0].metadata.orig_elements[0].metadata.last_modified remains None. If the serialized element has an older timestamp, the original retains that older value. Before this refactor, the partition function applied the explicit timestamp before its chunking decorator ran. This breaks provenance for consumers that inspect or serialize retained originals.

Suggested fix: apply a truthy explicit metadata_last_modified to rehydrated input elements before _apply_chunking(). Keep the filename attachment guard separate and preserve stored timestamps when no explicit override is supplied; do not substitute the JSON container's filesystem timestamp. The post-chunk assignment can remain for custom chunkers that create their own elements.

Add a regression covering serialized elements with missing and existing timestamps under basic and by_title. Assert the override on both chunks and retained originals, including after serialization, and preserve existing no-override behavior. The current originals test covers arbitrary JSON; the serialized timestamp tests do not request chunking.

(authored by codex)

@rohan-patnaik

Copy link
Copy Markdown
Author

Fixed the explicit-timestamp provenance regression in 3f89166: rehydrated originals receive the override before chunking, while no-override timestamps stay unchanged. The eight new basic/by_title and serialization cases reproduce four failures before the fix; all 72 JSON tests, repository Ruff and version-sync checks now pass. Also merged current main and preserved its releases, advancing this change to 0.27.17.

Signed-off-by: Rohan Patnaik <rohan-patnaik@users.noreply.github.com>
@rohan-patnaik

Copy link
Copy Markdown
Author

Merged current main into 8e34e6c. The conflict was limited to the changelog: upstream 0.27.17 is retained and this PR’s metadata entry/version is now 0.27.18. All 72 JSON tests, whole-tree Ruff, scoped formatting and version synchronization pass. GitHub confirms the exact head is mergeable; external workflow authorization remains with maintainers.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

bug(json): partition_json() does not preserve original element_id or metadata

2 participants