fix(json): preserve rehydrated element identity - #4504
rohan-patnaik wants to merge 4 commits into
Conversation
Review findingsI found three Medium-priority compatibility regressions introduced by this refactor. The core unchunked rehydration behavior—preserving serialized element IDs and stored metadata—looks correct. [Medium] Preserve
|
2b5617d to
9e42507
Compare
|
Addressed all three compatibility findings in
I also advanced the development version to |
Signed-off-by: Rohan Patnaik <rohan-patnaik@users.noreply.github.com>
|
Resolved the new release conflicts in 09d7669: preserved main’s changelog entries and advanced this change to 0.27.13. The JSON implementation is unchanged. All 64 JSON tests, Ruff lint/format, and version-sync validation pass; GitHub now reports the PR mergeable. |
There was a problem hiding this comment.
Reviewed head 09d7669e5d8c0b022d06796a89a8433070551206 against base 74e2fea06836169ee7749d61ce7d555ce4b83051, including the full production Oracle Pro review and independent source verification. The three earlier findings are fixed: arbitrary-JSON originals retain timestamps, attachment filenames survive an explicit filename override, and registered chunkers receive the complete bound arguments on all four dispatch paths.
[Medium] Apply explicit timestamp overrides before chunking rehydrated elements
Affected code: unstructured/partition/json.py:142-155.
The rehydration branch calls _apply_chunking() before applying metadata_last_modified. Built-in chunkers create separate chunk metadata and retain the input elements in metadata.orig_elements (unstructured/chunking/base.py:878-880,961-977; tables retain a separate copy at 1185-1188). Updating the returned chunk at line 155 therefore leaves its originals with missing or stale timestamps.
For example:
chunks = partition_json(
text='[{"type":"NarrativeText","element_id":"source-id",'
'"text":"Saved document content."}]',
metadata_last_modified="2020-07-05T09:24:28",
chunking_strategy="basic",
)The returned chunk receives the explicit timestamp, but chunks[0].metadata.orig_elements[0].metadata.last_modified remains None. If the serialized element has an older timestamp, the original retains that older value. Before this refactor, the partition function applied the explicit timestamp before its chunking decorator ran. This breaks provenance for consumers that inspect or serialize retained originals.
Suggested fix: apply a truthy explicit metadata_last_modified to rehydrated input elements before _apply_chunking(). Keep the filename attachment guard separate and preserve stored timestamps when no explicit override is supplied; do not substitute the JSON container's filesystem timestamp. The post-chunk assignment can remain for custom chunkers that create their own elements.
Add a regression covering serialized elements with missing and existing timestamps under basic and by_title. Assert the override on both chunks and retained originals, including after serialization, and preserve existing no-override behavior. The current originals test covers arbitrary JSON; the serialized timestamp tests do not request chunking.
(authored by codex)
|
Fixed the explicit-timestamp provenance regression in 3f89166: rehydrated originals receive the override before chunking, while no-override timestamps stay unchanged. The eight new basic/by_title and serialization cases reproduce four failures before the fix; all 72 JSON tests, repository Ruff and version-sync checks now pass. Also merged current main and preserved its releases, advancing this change to 0.27.17. |
Signed-off-by: Rohan Patnaik <rohan-patnaik@users.noreply.github.com>
|
Merged current main into |
Summary
Preserve serialized element IDs and source metadata when
partition_json()rehydrates Unstructured output. Generic JSON metadata and deterministic ID processing now apply only to arbitrary JSON, while explicit metadata overrides remain supported.Fixes #3365.
Verification
One existing deep-recursion expectation is excluded because it fails identically on current main under Python 3.13.