Skip to content

perf(chunking): bound image payload and original-element serialization memory (0.27.14) - #4517

Open
CyMule wants to merge 8 commits into
mainfrom
perf/bound-orig-element-serialization-memory
Open

CyMule wants to merge 8 commits into
mainfrom
perf/bound-orig-element-serialization-memory

Conversation

@CyMule

@CyMule CyMule commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

A stream of empty images can hold every base64 payload until a later text element even when include_orig_elements=False and those image fields are dropped from output. Fold metadata incrementally for contiguous empty-image runs and release excluded image payloads immediately, preserving metadata order and chunk boundaries.

Original-element serialization now adjusts and encodes one element at a time, escapes large JSON strings in bounded fragments (values without large strings still use one standard json.dumps call), and spills compressed buffers above 1 MiB to temporary disk. Avoid the redundant deep copy of the whole original-element graph in ElementMetadata.to_dict().

Proposed library release: 0.27.14, together with the other pending library memory changes.

Validation

  • Chunking, original-element staging and element suites: 764 passed, 24 skipped locally; 137 new regression cases also pass in the named SND runtime.
  • 120 mixed-image/text/title/table/page-break cases with overlap preserve chunk boundaries and metadata against the identity-preserving path. Original Image identity remains intact with the default include_orig_elements=True.
  • Exact legacy JSON field order is preserved, including continuation metadata and empty originals.
  • Exact legacy compressed bytes match for empty inputs, multiple elements, long Unicode/control-character strings, metadata precision and sorted JSON keys on the same zlib implementation.
  • Actual SND 128 streamed 4 MiB image payloads: chunking RSS 646 → 142 MiB, runtime 0.215 → 0.027 s; serialization RSS 1,674 → 146 MiB, runtime 3.770 → 3.217 s. Output hashes match exactly.
  • At a 384 MiB pod limit, both baseline cases are OOMKilled; both candidates pass with zero cgroup OOM events.
  • Repository Ruff checks/format and the actual Linux version-sync checker pass.

Operational impact

Compressed original-element buffers above 1 MiB use temporary disk. Final serialized output and metadata actually included in chunks still require memory proportional to output. include_orig_elements=True continues retaining original objects; changing that identity contract would require a separate API decision.

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All reported issues were addressed across 8 files

Shadow auto-approve: would not auto-approve because issues were found.

Re-trigger cubic

Comment thread scripts/performance/benchmark_chunking_memory.py Outdated

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 1 file (changes from recent commits).

Shadow auto-approve: would not auto-approve. This PR does not meet the repository auto-approval settings.

Re-trigger cubic

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 2 files (changes from recent commits).

Shadow auto-approve: would not auto-approve. This PR does not meet the repository auto-approval settings.

Re-trigger cubic

CyMule added 2 commits October 1, 2026 22:31
Values without a string longer than the fragment size are emitted with one
json.dumps call, so typical original elements no longer pay the per-node
generator cost. Large strings are still escaped in bounded fragments.

@cubic-dev-ai cubic-dev-ai Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 5 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="unstructured/staging/base.py">

<violation number="1" location="unstructured/staging/base.py:284">
P2: `_has_large_string` ignores dictionary keys, so a large key in valid metadata such as `data_source.record_locator` bypasses bounded fragmentation and allocates the entire JSON value in one fragment. Include keys in this traversal.</violation>
</file>

Shadow auto-approve: would not auto-approve because issues were found.
Tip: Review your code locally with the cubic CLI to iterate faster.

Re-trigger cubic

Comment thread unstructured/staging/base.py Outdated
@CyMule CyMule changed the title perf(chunking): bound image payload and original-element serialization memory (0.27.12) perf(chunking): bound image payload and original-element serialization memory (0.27.14) Oct 2, 2026

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0 issues found across 2 files (changes from recent commits).

Shadow auto-approve: would not auto-approve. This PR does not meet the repository auto-approval settings.

Re-trigger cubic

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 8 files

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="scripts/performance/benchmark_chunking_memory.py">

<violation number="1" location="scripts/performance/benchmark_chunking_memory.py:26">
P3: The `--case originals` benchmark never exercises the >1 MiB temp-disk spill path it is meant to validate. `image_base64` is `x` repeated, and that payload dominates the stream; runs of an identical byte compress to near nothing with `zlib.compressobj`, so the compressed output written to `SpooledTemporaryFile(max_size=1 MiB)` stays far below 1 MiB (128 images' metadata compresses to only tens of KB), and the spooled buffer never rolls over to disk. The peak-RSS numbers and PR claim about the disk-spill optimization are therefore untested for the large-payload case that motivated them. Use an incompressible payload (e.g. base64 of random bytes) so the compressed buffer actually exceeds the 1 MiB roolover threshold.</violation>
</file>

Shadow auto-approve: would not auto-approve because issues were found.

Re-trigger cubic

Comment on lines +26 to +27
metadata=ElementMetadata(
image_base64="x" * (args.payload_mib * 1024 * 1024),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: The --case originals benchmark never exercises the >1 MiB temp-disk spill path it is meant to validate. image_base64 is x repeated, and that payload dominates the stream; runs of an identical byte compress to near nothing with zlib.compressobj, so the compressed output written to SpooledTemporaryFile(max_size=1 MiB) stays far below 1 MiB (128 images' metadata compresses to only tens of KB), and the spooled buffer never rolls over to disk. The peak-RSS numbers and PR claim about the disk-spill optimization are therefore untested for the large-payload case that motivated them. Use an incompressible payload (e.g. base64 of random bytes) so the compressed buffer actually exceeds the 1 MiB roolover threshold.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At scripts/performance/benchmark_chunking_memory.py, line 26:

<comment>The `--case originals` benchmark never exercises the >1 MiB temp-disk spill path it is meant to validate. `image_base64` is `x` repeated, and that payload dominates the stream; runs of an identical byte compress to near nothing with `zlib.compressobj`, so the compressed output written to `SpooledTemporaryFile(max_size=1 MiB)` stays far below 1 MiB (128 images' metadata compresses to only tens of KB), and the spooled buffer never rolls over to disk. The peak-RSS numbers and PR claim about the disk-spill optimization are therefore untested for the large-payload case that motivated them. Use an incompressible payload (e.g. base64 of random bytes) so the compressed buffer actually exceeds the 1 MiB roolover threshold.</comment>

<file context>
@@ -0,0 +1,62 @@
+        yield Image(
+            text="",
+            element_id=f"image-{index}",
+            metadata=ElementMetadata(
+                image_base64="x" * (args.payload_mib * 1024 * 1024),
+                filename="streamed.pdf",
</file context>
Suggested change
metadata=ElementMetadata(
image_base64="x" * (args.payload_mib * 1024 * 1024),
image_base64=base64.b64encode(os.urandom(args.payload_mib * 1024 * 1024)).decode(),

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant