Skip to content

fix(docx-compare): align identical text across run re-segmentation - #1148

Merged
stevenobiajulu merged 8 commits into
mainfrom
agent/issue-1142
Oct 4, 2026
Merged

stevenobiajulu merged 8 commits into
mainfrom
agent/issue-1142

Conversation

@stevenobiajulu

@stevenobiajulu stevenobiajulu commented Oct 3, 2026 •

Copy link
Copy Markdown
Member

Closes #1142

Problem

refineSimpleRunGap only tokenized concatenated text when every run in the gap shared one run-property signature. Any mixed-formatting paragraph fell back to tokenizing each run on its own, so identical text split into runs differently () + , in two runs against ), in one) produced different token sequences and spurious w:del/w:ins pairs. A safe-docx save of a one-word edit re-segments the NVCA Voting Agreement preamble (18 runs to 5), and the redline showed DEL ) DEL , INS ), twice and DEL ) DEL . INS ). once next to the real edit.

Separately, the tagged-token-v1 atom statistics tokenized each w:t independently, so a re-segmentation-only change reported 0/0 ranges but 3/6 inserted/deleted atoms.

Fix

  • Serializer (taggedTreeSerializer.ts): tokenizedRuns now tokenizes each maximal group of adjacent runs with the same run-property signature and the same operation provenance as one concatenated string. Token boundaries are forced where either changes, so no token mixes formatting or revisionAttributionRanges operations. Tokens keep the paragraph-level offset (start) and the run that owns their first character (run), the same contract the old concatenated path had, so emitCommonToken (which splits a common token at every run boundary on both sides and emits per-fragment emitCommonRun property deltas) and the runFragment/provenance lookup for deleted/inserted tokens are reused unchanged. When the whole gap has one signature, the output is byte-identical to the old concatenated path.
  • Readable bridges (docx-markdoc: Markdoc builds still need raw-OOXML post-processing (tables, side stories, greenfield, formatting, readable redlines) #998): bridgeMatches still requires one signature across the whole gap. This is deliberate. The docx-markdoc: Markdoc builds still need raw-OOXML post-processing (tables, side stories, greenfield, formatting, readable redlines) #998 contract says readable grouping leaves formatting boundaries intact. Applying it per signature group would let a deletion/insertion chain coalesce across a bold/plain boundary (see the new keeps readable whitespace bridges to gaps with one formatting signature test), and that would be a policy extension, not a bug fix. Single-signature gaps behave exactly as before. The line itself is unchanged, which keeps the diff small near the docx-markdoc: Markdoc builds still need raw-OOXML post-processing (tables, side stories, greenfield, formatting, readable redlines) #998 code.
  • Atom stats (taggedTreeShadow.ts): comparisonAtomKeys now treats adjacent w:t text within a paragraph as one stream, so the keys depend only on the text and not on run boundaries. Any other leaf (tab, br, field chars, delText, …) and every paragraph boundary end the stream. The atom metric does not carry formatting (that is formatChangeAtoms), so it is segmentation-invariant rather than signature-grouped. When a paragraph exists on both sides, its alignment (unalignedParagraphAtoms) compares whitespace one character per key, which keeps the spaces around a removed or moved run (prefix + Company + suffix against prefix suffix) from turning into a different token. Any word or control match outweighs every whitespace match, so spaces never displace an unchanged word. The unaligned characters of one whitespace token weigh one atom together. A match involving moved content weighs half an unmoved one, so the existing subtractMoves subtraction stays exact. Whole-element weights use the plain concatenated tokens. One intended difference from main: shortening a run of spaces counts as one deleted atom, not one deleted plus one inserted.

Results

NVCA Voting Agreement, orig vs the safe-docx-saved Company→SMOKEWORD revision:

insertions/deletions atoms ins/del
before 4 / 7 4 / 7
after 1 / 1 1 / 1
merge-only control (before → after) 0/0 → 0/0 3/6 → 0/0

The redline for that paragraph is now only DEL Company / INS SMOKEWORD.

Strategy-differential manifest changes (reviewed, all reductions in spurious churn)

  • checked-in/ILPA: ranges 844/758 → 821/740. All 10 paragraphs whose redline changed lose re-segmentation churn (for example DEL Act INS Ac INS t → nothing, and DEL If any DEL Limited INS If INS any INS Limit INS ed → DEL If any INS If any). Atoms 3662/2101 → 2932/1385.
  • Atom-only reductions where words were split across runs: split-run-boundary-change (t + he was 2 atoms, now 1), p-unit-agreement-v2, and the NVCA paragraph-deletion rows for COI, Indemnification, MRL and SPA.
  • Reject/accept projection safety checks in the harness pass for every row.

Tests

  • taggedTreeSerializer.test.ts: new mixed-format run re-segmentation (#1142) block with minimal XML reproductions, run in both directions: a one-word edit with )+, vs ), (exactly one del and one ins, plus stats), re-segmentation alone (0 ranges and 0 atoms), a formatting change at a run boundary next to a text edit (w:rPrChange still emitted), a formatting boundary that splits a word (still a token boundary), the readable-bridge gate, adjacent runs attributed to different operations (separate revisions, attribution resolves), atom weights around a removed word run, a pure inline move, a move plus a deletion, and a space deleted next to a moved run, plus whitespace edits (a moved space run, the same spaces split differently by run boundaries, and a deleted paragraph's weight under two segmentations). The first three fail on main. The attribution and atom tests cover regressions that review found in this PR's earlier commits.
  • Real corpus (real-corpus-paragraph-deletion.test.ts): a new real-corpus run re-segmentation block derives the save-style re-segmentation of the Voting Agreement preamble (merging adjacent same-rPr plain text runs) with and without the one-word edit. It asserts 1/1 and 0/0 ranges and atoms. Both fail on main.
  • I ran npm run test:real-corpus -w @usejunior/docx-compare locally against the SHA-verified corpus (scripts/prepare_real_comparison_corpus.mjs, *_REQUIRED=1): 30/30 passed.
  • npm run build, lint:workspaces, check:allure-labels, check:allure-quality, check:allure-filenames, and the test:run suites for docx-compare, docx-core, docx-mcp and docx-markdoc all pass.

Known gaps (not changed here)

@vercel

vercel Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
site Ready Ready Preview Oct 4, 2026 1:46am UTC

Request Review

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-03T23:55:27.992159Z c002df6 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c002df6914

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +129 to +130
if (current.localName === 't') {
text += current.textContent ?? '';

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Version the changed atom weighting

For paragraphs where a word is split across multiple w:t leaves, concatenating the leaves before tokenization changes the public metric unit—for example, t + he was two tagged-token-v1 atoms and is now one—while deriveTaggedTreeStats still reports atomMetricVersion: 'tagged-token-v1'. This makes persisted metrics and consumer thresholds silently incomparable and violates the active docx-comparison requirement that a weighting change use a new version value; either retain the v1 weighting or introduce and publish a new metric version.

AGENTS.md reference: openspec/AGENTS.md:L456-L456

Useful? React with 👍 / 👎.

Comment on lines +1444 to +1445
for (const value of tokenizeComparisonText(group.map(runText).join(''))) {
tokens.push({ value, run: ownerAt(offset), start: offset });

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve attribution boundaries when grouping runs

When adjacent same-rPr runs in a mixed-format replacement gap belong to separate revisionAttributionRanges but together form one token, this assigns the whole token only to the run owning its first character. The later insertion/deletion path reads provenance solely from that run, so an operation attached to a subsequent run emits no attributed revision and resolveTaggedRevisionAttributions throws operation ... has no emitted attributed revision; group on operation provenance as well, or split side-only tokens at source-run boundaries before wrapping them.

Useful? React with 👍 / 👎.

@usejunior-llm-gate

usejunior-llm-gate Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

LLM gate (advisory)

All evaluated rules passed - 5 pass, 0 warn, 0 error, 11 skipped, 16 total

Findings

None.

All 16 rules (5 evaluated, 11 skipped)
Rule Verdict Detail
read_file response metadata parity SKIPPED paths not touched by this PR
Live DOM namespace-safe OOXML writes SKIPPED paths not touched by this PR
Complex-field revisions preserve complete accept/reject state machines PASS The PR does not touch field atomization, validateFieldStructure, w:fldChar, w:instrText, w:delInstrText, or collapsed-field comparison logic, focusing instead on run re-segmentation and paragraph atom alignment.
Field validation per story, not global SKIPPED paths not touched by this PR
Revision IDs seeded from all revision-bearing side parts SKIPPED paths not touched by this PR
Accept/reject sweep side parts and caches PASS The precondition is not met as this PR does not touch DocxDocument.acceptChanges, DocxDocument.rejectChanges, REVISION_STORY_PART_PATHS, accept_changes, reject_changes, or side-part revision markup, focusing instead on comparison run re-segmentation and alignment stats (#1142).
DocumentViewNode.heading stays canonical SKIPPED paths not touched by this PR
AI-author parity across entry points SKIPPED paths not touched by this PR
Property-change wrapper discipline SKIPPED paths not touched by this PR
SUPPORT.md Table A drift vs. implementation SKIPPED paths not touched by this PR
Table A / Table B boundary on side-part revisions SKIPPED paths not touched by this PR
Canonical-emission surface completeness SKIPPED paths not touched by this PR
Unit-test quality (avoid tautological / change-detector tests) PASS The test assertions in packages/docx-compare/src/integration/real-corpus-fixtures.test.ts:54, packages/docx-compare/src/integration/real-corpus-paragraph-deletion.test.ts:269, and packages/docx-compare/src/tagged/taggedTreeSerializer.test.ts:1072 are high-quality, independent of the SUT, and constructed from first principles. Expected values are explicit literals verifying precise stats and text, avoiding SUT mocking or snapshotting, and directly exercising the run re-segmentation bug/feature (#1142).
Re-derived facts vs canonical sources PASS The PR does not add logic re-computing facts canonical to other parts of the codebase (such as fldChar walks or footnote references). The changes in packages/docx-compare/src/tagged/taggedTreeSerializer.ts:1425 and taggedTreeShadow.ts:134 refine the comparison alignment logic directly without duplicating or re-implementing existing helper functions.
.openspec tag ↔ test-assertion drift PASS The PR does not add, move, or change any .openspec tags on any test or it blocks in the modified files, so the precondition is not met.
Library stays general (no downstream-domain leakage) SKIPPED paths not touched by this PR

@codecov

codecov Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

@stevenobiajulu

Copy link
Copy Markdown
Member Author

Peer review, round 1 (Codex, codex-review registry default)

Verdict: CHANGES_REQUIRED. The reviewer did not write a formal verdict line before exiting, but its execution trace reproduced two regressions against the pre-change build, and I confirmed both.

Adjudication

  1. Blocking, accepted: operation attribution. Two adjacent same-format revised runs attributed to different operations (blue → A, stone → B) were concatenated into one token (bluestone). Publication then threw operation B has no emitted attributed revision. Main emits two attributed insertions. Fix: the tokenization group key is now the run-property signature plus the operation provenance, so a token never spans two operations. Regression test: keeps adjacent runs from different attributed operations in separate revisions (checks the per-operation w:ins wrappers and resolveTaggedRevisionAttributions).
  2. Blocking, accepted: atom counts around moves. Concatenating paragraph text merged the spaces on either side of a moved run (prefix + cat + suffix against prefix suffix). A pure inline move reported 3/3 atoms (0/0 on main), and the same artifact hit an ordinary deletion of a word that sits in its own run. Fix: moved subtrees keep their own token boundaries, so the existing subtractMoves subtraction is exact. Paragraph alignment compares whitespace per character, while each run of unaligned whitespace still weighs one atom. I first tried excluding moved subtrees from the streams instead. That inflated ILPA atoms to 7564/7336, because tagged moves can sit in paragraphs whose two sides are otherwise identical, and the clamp in subtractMoves is what makes those safe. So I kept the subtraction. Regression test: weighs atoms without the whitespace merged around removed or moved runs (removed word run 1/0, pure move 0/0, move plus deletion 0/1, matching main).

Other checks from the trace that hold up: the full docx-compare suite passed (815 tests); the real-corpus suite passed against the verified cache; the fixture helper reduces the Voting Agreement preamble from 18 runs to 5 and changes only Company→SMOKEWORD; the old build reproduces 4/7 on that fixture.

Fixes are in 3c6d9bf. Manifest regenerated: ILPA atoms are now 2926/1385 (main 3662/2101), and ranges are unchanged from round 1 (821/740). A confirmation round follows.

@stevenobiajulu

Copy link
Copy Markdown
Member Author

Peer review, round 2 (confirmation, Codex)

Verdict: CHANGES_REQUIRED. The round-1 fixes were confirmed: A/B attribution resolves, a pure move is 0/0, and move plus replacement and move plus deletion match main. The full package suite (817), the tagged suite and the real-corpus suite (30/30) passed. The reviewer found two new regressions caused by my round-1 whitespace change.

Adjudication

  1. Blocking, accepted: splitting whitespace into per-character keys let a long space run outweigh words in the LCS. a b c d e → a b c d e reported 5/4 atoms (main 2/1), and the 20-word variant reported 21/20. I confirmed this, and my first patch (a weighted LCS in which word matches dominate) still differed from main on whitespace-only edits (a b c → a b c gave 0/2 against main's 2/2). Fix: whitespace stays whole tokens, as on main, split only where a w:t boundary falls inside one. That is enough to keep the spaces around a removed or moved run as two tokens. Word and punctuation tokens still span runs, which is the docx-compare: run re-segmentation produces spurious punctuation del/ins in mixed-format paragraphs #1142 fix.
  2. Blocking, accepted: with a space deleted next to a moved cat run, the alignment could match the moved run's spaces instead of the independent one, and the move subtraction then erased the real deletion (0/0, expected 0/1). Fix: a match involving moved content weighs half an unmoved match, so moved keys stay unaligned and are subtracted at exactly their standalone weight.

Also added: an unaligned gap whose two sides spell the same unmoved text (for example Keep the + double space. against Keep the double space.) counts as unchanged. Without it, segmentation inside whitespace would still produce atoms.

Every case in the reviewer's probes now matches main: both of its regression assertions pass; the whitespace deletions, the replacement and the space-dominance probes give 2/2, 1/1 and 2/2; the hyperlink, tab, break and space moves give 0/0. The Voting Agreement fixture stays at 1/1 and the merge-only control at 0/0. Unit tests cover all of these. The manifest is unchanged from round 2 (ILPA atoms 2926/1385). Fixes are in 4308a19, and a confirmation round follows.

A mixed-formatting paragraph tokenized each run on its own, so the same
text split into runs differently (`)` + `,` against `),`) produced
spurious deletion/insertion pairs around unchanged punctuation. A safe-docx
save of a one-word edit in the NVCA Voting Agreement preamble reported
7 deletions and 4 insertions instead of 1 and 1.

Tokenize each maximal group of adjacent runs that share a run-property
signature as one text stream, reusing the concatenated-path token/offset
machinery; a formatting change still forces a token boundary. Readable
whitespace bridges (#998) stay limited to single-signature gaps so they
never coalesce across a formatting change.

The tagged-token-v1 atom statistics tokenized each w:t on its own and
disagreed with the range counts (3/6 atoms for a 0/0-range
re-segmentation). Atom keys now tokenize adjacent w:t text as one stream
per paragraph. The strategy-differential manifest is updated: ILPA loses
23/18 spurious insertion/deletion ranges, and atom counts drop where
words were split across runs.

Closes #1142
Peer review of the run re-segmentation fix found two regressions:

- Adjacent same-format runs attributed to different operations were
  concatenated into one token, so publication failed with "operation B
  has no emitted attributed revision". Tokenization groups now also break
  where operation provenance changes.
- Paragraph atom alignment over one concatenated text stream merged the
  spaces on either side of moved or removed runs (`prefix ` + `cat` +
  ` suffix` against `prefix  suffix`), so a pure inline move reported 3/3
  atoms and deleting a word run reported extra whitespace atoms. Moved
  subtrees now keep their own token boundaries before their standalone
  weight is subtracted, and paragraph alignment compares whitespace per
  character while still weighing each run of unaligned whitespace as one
  atom.

Ref: #1142
The confirmation review found that aligning whitespace one character at
a time let a long space run outweigh unchanged words (moving ten spaces
across a sentence reported 5/4 atoms instead of 2/1), and that a space
deleted next to a moved run could be absorbed when the moved run's
weight was subtracted.

Paragraph atom streams now keep whitespace as whole tokens, split only
where a w:t boundary falls inside one, so the spaces on either side of a
removed or moved run stay two tokens; word and punctuation tokens still
span runs. Alignment weighs matches involving moved content at half an
unmoved match, so moved keys stay unaligned and are subtracted exactly.
An unaligned gap whose two sides spell the same unmoved text (a double
space split by a run boundary on one side only) counts as unchanged.

Ref: #1142
The round-3 review showed that splitting whitespace tokens at w:t
boundaries still left atoms when the same spaces were split differently
(`a ` + `  b` against `a  ` + ` b`), and that a deleted paragraph's
weight depended on its run boundaries.

Whitespace in a paragraph stream is now one key per character of each
concatenated whitespace token, so keys depend only on the text. Word and
control matches outweigh every whitespace match, so spaces never displace
an unchanged word, and the unaligned characters of one whitespace token
weigh one atom together. Whole-element weights use the plain concatenated
tokens. A shortened space run now counts as one deletion rather than a
deletion plus an insertion.

Ref: #1142
@stevenobiajulu

Copy link
Copy Markdown
Member Author

Peer review, round 3 (confirmation, Codex)

Verdict: CHANGES_REQUIRED. Both round-2 regressions were confirmed fixed. The full package suite (818), the tagged suite and the real-corpus suite (30/30) passed, and a 1560-pair sweep of the gap-cancellation rule found no hidden text change.

Adjudication

  1. Blocking, accepted: run-segmentation invariance was incomplete for whitespace. Splitting whitespace tokens at w:t boundaries meant a + b against a + b (the same three spaces) still produced 1/1 atoms (24 of 140 cases in the reviewer's sweep), and a deleted paragraph weighed 3 or 4 atoms depending on its runs. Main behaves the same way, but this PR claims segmentation invariance, so I agree it blocks. Fix (e32bde1): whitespace in a paragraph stream is one key per character of each concatenated whitespace token, so the keys depend only on the text. Any word or control match outweighs every whitespace match, so the round-2 word-displacement case stays fixed. The unaligned characters of one whitespace token weigh one atom. Whole-element weights use the plain concatenated tokens. The reviewer's sweep now reports 140 cases, 0 nonzero results, and its two round-3 regression assertions pass.
    • Intended semantic change: with whitespace now compared per character, shortening a space run counts as one deletion. So a b c → a b c gives 0/2 atoms (main 2/2), and the round-2 long-space-run probe gives 1/1 (main 2/1). Main's whole-token rule cannot be segmentation-invariant: the three-spaces case above already contradicts it. The CHANGELOG entry states the change.
  2. Non-blocking, accepted as a follow-up: a moved key can absorb an identical real edit (a hyperlink cat moves away while dog becomes cat gives 0/0 atoms instead of 1/1). The same happens on main. Making moved keys unmatchable fixes it but changes ILPA counts a lot, because tagged moves sit in paragraphs that are otherwise identical. Filed as docx-compare: moved content can absorb a real edit in insertedAtoms/deletedAtoms #1153.

I rebased onto origin/main (CHANGELOG conflict only) and reran the full gate on the rebased head 02e2417. Build, lint, the allure checks, the docx-compare, docx-core, docx-mcp and docx-markdoc suites, and the real corpus (30/30) all pass. Manifest: ILPA atoms 2932/1385; ranges unchanged at 821/740. A confirmation round follows.

@stevenobiajulu

stevenobiajulu commented Oct 4, 2026 •

Copy link
Copy Markdown
Member Author

Peer review, round 4 (confirmation, Codex)

Verdict: CHANGES_REQUIRED. The round-3 blocking cases are confirmed fixed (boundary shift 0/0, paragraph-deletion weight 3/3). The reviewer accepted the per-character whitespace metric ("[1,1] is reasonable under the documented whitespace metric"; 300 whitespace-only cases with repeated words, 0 mismatches). The full suite (818), the tagged suite and the real corpus (30/30) passed.

Adjudication

  1. Blocking finding: rebutted as out of scope, filed as docx-compare: an unchanged run matched at a different text offset splits a re-segmented gap into spurious del/ins #1154, and the claim narrowed. The new case (bold X, then [" ", "[", " n"] against [" [", " ", "n"]) gives 1/2 ranges with 0/0 atoms. I reproduced it on the pre-change build with the same 1/2 result, so this PR did not introduce it. The mechanism is not docx-compare: run re-segmentation produces spurious punctuation del/ins in mixed-format paragraphs #1142's. docx-compare: run re-segmentation produces spurious punctuation del/ins in mixed-format paragraphs #1142's root cause is that tokenizedRuns tokenized each run on its own within a refined gap. Here, tagged-tree construction matches the single-space run as a both node at a different text offset on each side, and emitNode's gap collection stops at that both node, so each half is refined separately. Fixing it means letting gap refinement step over matched both runs, or adding a paragraph-level per-character equality pass. Both change output well beyond re-segmentation and need their own manifest review. The reviewer itself labelled it "a remaining instance of docx-compare: run re-segmentation produces spurious punctuation del/ins in mixed-format paragraphs #1142's symptom, rather than a round-4 regression".
  2. The docx-compare: moved content can absorb a real edit in insertedAtoms/deletedAtoms #1153 move/atom item is unchanged and tracked there.

A confirmation round follows to adjudicate the scoping.

@stevenobiajulu

Copy link
Copy Markdown
Member Author

Peer review, round 5 (confirmation, Codex)

Verdict: APPROVE. "Blocking: None. Your scoping argument holds… I withdraw the round-4 blocker under the revised scope."

The reviewer verified this by running both builds: the #1154 repro gives identical statistics on the merge base and on HEAD (1/2 ranges, 0/0 atoms); the NVCA Voting Agreement fixture goes from 4/7 ranges and atoms on the base to 1/1 on HEAD; shortening a space run counts one deleted atom. The full docx-compare suite (818), the real-corpus suites including both new NVCA cases and the manifest, workspace lint and the conformance checks all passed.

One non-blocking nit, applied: the CHANGELOG sentence "a re-segmentation-only change now reports zero atoms as well as zero ranges" overstated the range guarantee. It now reads "zero atoms, and zero ranges unless an unchanged run is matched at a different text offset (#1154)" (CHANGELOG only, in the head commit after 231f3ea).

Review summary across rounds: round 1 found an operation-attribution regression and move-atom miscounts (fixed); round 2 found whitespace dominance and a move-subtraction edge (fixed); round 3 found incomplete whitespace segmentation invariance (fixed with per-character whitespace keys, which changes how whitespace-only edits are weighted; accepted by the reviewer); round 4 raised a pre-existing tree-matching case (scoped out to #1154 and the claims narrowed); round 5 approved. Follow-ups: #1149, #1150, #1153, #1154.

The real-corpus suites are excluded from the default run, so the new
fixture helper counted as uncovered and dropped docx-compare below its
coverage ratchet. Exercise it on a synthetic preamble-shaped package:
adjacent same-format plain runs merge across punctuation, a tab run and
bookmark boundaries are respected, only the target word changes, and the
comparison reports 1/1 for the edit and 0/0 for re-segmentation alone.

Ref: #1142
@stevenobiajulu
stevenobiajulu merged commit aca5f2b into main Oct 4, 2026
27 checks passed
@stevenobiajulu
stevenobiajulu deleted the agent/issue-1142 branch October 4, 2026 02:12
@stevenobiajulu

Copy link
Copy Markdown
Member Author

Post-merge smoke (origin/main @ aca5f2b)

I built aca5f2be, the squash of this PR on main, from a clean detached checkout (npm install && npm run build). Then I ran compareDocuments(orig, rev) on real NVCA forms. Each rev.docx is the original after one safe-docx MCP replace_text (Company → SMOKEWORD) and save.

Document ins / del ranges ins / del atoms Redline content
NVCA Voting Agreement 1 / 1 1 / 1 DEL Company, INS SMOKEWORD; no ), / ). churn (was 4 / 7 before this PR)
NVCA Investors' Rights Agreement 1 / 1 1 / 1 DEL Company, INS SMOKEWORD
NVCA Stock Purchase Agreement 1 / 2 1 / 1 DEL Company, INS SMOKEWORD, plus one empty w:del

The extra SPA deletion range is an empty run inside a REF field that the safe-docx save dropped. Main produced the same output before this PR; it is tracked in #1150 and does not come from this change.

Result: PASS. The Voting Agreement case from #1142 now yields one deletion and one insertion, and no punctuation is deleted and re-inserted. Vercel status on aca5f2be is success.

This branch was successfully deployed

1 active deployment
Preview — d2f6fb34 Deployed Oct 4, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fix PR type: bug fix

Projects

None yet

Development

Successfully merging this pull request may close these issues.

docx-compare: run re-segmentation produces spurious punctuation del/ins in mixed-format paragraphs

1 participant