Skip to content

ml: split comma-joined clauses in seeds; lint the whole lecture (round 4) - #331

Merged
mmcky merged 2 commits into
mainfrom
ml/splice-lint-repair
Sep 28, 2026
Merged

mmcky merged 2 commits into
mainfrom
ml/splice-lint-repair

Conversation

@mmcky

@mmcky mmcky commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This PR updates the Malayalam calibration tooling from the editor's fourth review round (QuantEcon/lecture-python-programming.ml#23, numpy, 73 suggestion blocks, merged 2026-09-28). Two changes, both in experiments/ml-benchmark/scripts/:

  1. ml_repair.py gains a fourth deterministic repair: it splits two finite clauses joined by a comma and a resumptive pronoun or connective into two sentences (… methods ഉണ്ട്, ഇവയെല്ലാം … → … methods ഉണ്ട്. ഇവയെല്ലാം …).
  2. ml_metrics.py now sees the whole lecture, and its line-level lints stop firing on hard wraps.

Nothing in src/ or dist-action/ changes, and CI does not cover these scripts. This is the prototype stage of #260 (post-processing), feeding #301 (lint gaps).

Why

The largest class in the round was comma-joined clauses where the editor writes a full stop: 20 flags, 23 comma joins removed.

Splice sites per 100 Malayalam prose lines
Nine numpy draws (v0.29.0–v0.29.2) 13.1–21.0 (mean 17.2)
His reviewed pages, rounds 2–4 (functions, matplotlib, numpy) 1.6–2.8
His round-1 page (python_by_example), before he was splitting these 8.7

Rule 12's "split … at comma splices" has been in the prompt throughout. Most sites don't come from an English comma splice: the model renders an English relative clause, so that, a participle or an appositive as comma + resumptive pronoun. A prompt rule does not reach this, and the programme's measurements say "always" rules land about a third of the time anyway, so this goes to lint + deterministic repair.

What changes

The repair (comma_splice_resumptive lint → ml_repair.py)

  • Match: a finite verb ending, then , , then a word from an explicit list — ഇത്, ഇതിനെ, അത്, അതിനെ, ഇവ/അവ with enumerated case endings, അതിനാൽ, അതേസമയം — with the clause carrying on after it on the same line.
  • Change: the comma becomes a full stop. No comma is ever added after the connective.
  • Guards:
    • the finite token must not open its sentence;
    • എല്ലാം ("all", which ends like the modal -ാം) and the fronted ശ്രദ്ധിക്കുക/ഓർക്കുക that rules 13/14 require are skipped;
    • a pronoun that ends the line (it opens a list) is not split, since that would leave a verbless അത്:;
    • the neither-nor -ഉം അല്ല, and a *അല്ല*, ഇത് … contrast are skipped;
    • a trailing "as shown" tag and a list-final , ഒപ്പം are skipped;
    • അങ്ങനെ, അതുകൊണ്ട് and പക്ഷേ are deliberately excluded: he writes , അങ്ങനെ himself.
  • One classifier: ml_repair calls ml_metrics.split_resumptive_splices, so the lint and the repair cannot disagree.

The report-only watch

  • comma_splice_watch lists every comma-joined finite clause, and comma_splice_rate gives the per-100-line figure.
  • His reviewed pages keep one to four per lecture, so zero is not the target. Use it as a secondary best-of-N key at most.

Prose scanner fixes (prose_line_info, with prose_lines kept as a wrapper)

  • A paragraph opening with ( was skipped as if it were a (label)= line. Five numpy lines were invisible to every lint, including a bracketed paragraph with no stop that is bare in 9/9 draws. Only whole-line labels are skipped now.
  • {note} and other prose-directive bodies are scanned (directive names case-insensitively). One of the round's splices sat in a {note}. Exercise-family directives, figures and code are never scanned. Exercise content is English by decision (D-2026-09-03-ml-all-exercise-content-stays-english, enforced by src/verbatim-directives.ts). As a consequence, python_by_example's round-1 exercise text, translated before that decision, is no longer linted; one hit is lost there (seed line 561, which he fixed).
  • Fences follow CommonMark. The old toggle treated any line starting with three backticks as a closer. It lost parity at a nested {hint} inside an unclosed {exercise-start} in python_by_example, then linted a code cell as prose and skipped the prose after it. A fence body now ends only at a bare closer at least as long as its opener.
  • Terminal punctuation is checked at paragraph ends, outside list items (wrapped continuations included), and capitalisation at paragraph starts. This removes the hard-wrap false positives: 11 fewer terminal hits and 3 fewer lowercase hits on the round-1 seed.
  • A closing bracket counts as terminated only after a stop, in the lint and the colon repair alike. (….) and ...) pass; (… memory) is flagged. He added that stop at ml#23 line 197 and in round 2, and accepted ...). The deep-copy line before a cell now gets his accepted ): (ml#23 line 920); it was bare in 7 of 9 draws and the repair used to skip it. A wholly bracketed paragraph is left to the lint.
  • A comma counts as an ending only before a list it opens, not before a code or math cell. That case is flagged and left to the lint, since the repair would write ,:. It occurs in no ml corpus. This follows Copilot's review.

Retired: the future-hortative watch. Since rule 13 landed it has made two hits on reviewed seeds and the editor acted on neither. It flagged a later-lecture promise he kept, and missed another.

Validation

Repair, round-4 seed:

  • It splits exactly the 13 comma-plus-pronoun sites, all on lines he flagged.
  • His text has the same full stop at 10; the other 3 he folded into a converb.
  • It reproduces three of his lines byte for byte (ml#23 lines 336, 494 and 841, used as fixtures).
  • It fires on 0 of the 70 lines he left untouched.

Repair, earlier rounds:

  • In the round-1 and round-2 seeds he removed the comma at all 12 such sites he edited: a full stop at 9, an em-dash at 2, a semicolon at 1. A 13th sits in a solution he reverted to English. Across all three rounds the full stop is his form at 19 of the 25 edited sites.
  • It makes no change to any of his four reviewed pages (python_by_example, functions, matplotlib, numpy).
  • On the two unsent current-engine draws it fires 12 and 15 times, all in-class on inspection.

Lint changes on his reviewed pages: hits are only removed, apart from two new true positives on python_by_example lines the old parser never saw: a lowercase short opener, and thought-ഉം planning-ഉം, which is the pair the editor's "always" answer on ml#22 covers.

Coverage diff: 38 files (the ml edition, every ml arm draw, the review seeds). Every added prose line is a ( paragraph, a prose-directive body, or prose the old toggle had lost after a parity error. Every Malayalam line no longer scanned sits inside an exercise-family directive, and those stay English by decision record.

Tests: python3 -m unittest discover -s experiments/ml-benchmark/scripts -p 'test_*.py' runs 16 stdlib tests. They cover his exact lines, one fixture per guard (including his reviewed python_by_example 275, which only the എല്ലാം guard protects), the parser cases (the {exercise-start}/{hint} desync, a code cell nested in a longer-fenced {Note}) and the bracket rule. Each of twelve single-guard mutations (removing a guard or reverting a regex) fails the suite.

Adversarial review before opening. Four reviewers ran the change over every ml corpus: parser, splice precision, lint regressions, and claims. Each finding went to two sceptics who reproduced it. 15 findings; 8 were confirmed by both, all low severity once verified, and all 8 are addressed in this PR:

  • a verbless split at a pronoun that ends a line (2 of 285 sites, both in the oldest draw);
  • the colon repair disagreeing with the new bracket rule;
  • wrapped list-item continuations flagged as bare endings;
  • overstated rounds-1–2 and reference-band figures;
  • tests that did not reach the guards.

None of the 15 findings showed the repair touching his reviewed text. Of the rejected findings, the exercise-body scope is kept on purpose (see above). The rest were either intended behaviour or inputs that do not occur in any corpus.

Follow-ups (not in this PR)

  • A repair for bracketed paragraphs with no inner stop, and a generalised -ഉം comma (negated pairs, multi-word items, chains).
  • The P2 lexical repairs from the round:
    • imperative see → നോക്കുക
    • adverbial element-wise → ഓരോ element-ലും
    • have already seen → കണ്ടുകഴിഞ്ഞു
    • standalone മുൻ → മുൻപത്തെ
  • Promotion into the engine under W2 — localisation becomes deterministic (after D1) [P1] #260, at model output only (file-processor.ts:689), never in the document-wide typography pass at :559. That pass would touch editor-reviewed sections.

The full per-flag analysis is in QuantEcon/project-translation reports/2026-09-28-ml-numpy-review-disposition.md.

🤖 Generated with Claude Code

…d 4)

From the fourth native review round (lecture-python-programming.ml#23,
numpy): the editor's largest class is two finite clauses joined by a comma
where he writes a full stop. Nine numpy draws run at 13-21 sites per 100
Malayalam prose lines; his reviewed pages at 1.6-2.8 since round 2.

- ml_metrics: comma_splice_watch, comma_splice_rate, and the repairable
  subset comma_splice_resumptive (finite verb + comma + a resumptive
  pronoun or അതിനാൽ/അതേസമയം, guarded against the forms he keeps).
- ml_repair: a fourth repair splits those sites with a full stop, using
  ml_metrics' classifier. Round-4 seed: 13 sites, all on lines he flagged,
  none on his 70 untouched lines; no change to his four reviewed pages.
- The prose scanner follows CommonMark fences, scans "(" paragraphs and
  prose-directive bodies, and checks punctuation at paragraph ends outside
  list items and capitalisation at paragraph starts. A closing bracket is
  terminated only after a stop, in the lint and the colon repair alike.
- The future-hortative watch is retired.
- 15 stdlib fixture tests from his own lines (test_ml_lints.py).

Experiment tooling only; nothing in src/ or dist-action/ changes.
Refs #260, #301.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Copilot AI lite review requested due to automatic review settings September 28, 2026 03:44

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The terminal-punctuation lint currently treats a trailing comma as acceptable even before code-cell/math blocks, which contradicts the documented convention and can suppress intended linting.

Review effort: Lite
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

This PR updates the Malayalam benchmarking/calibration scripts under experiments/ml-benchmark/scripts/ to (1) add a deterministic “comma splice + resumptive” sentence-splitting repair and (2) improve prose scanning so linting operates over the full lecture while avoiding hard-wrap false positives. It also adds fixture-driven stdlib tests for these behaviors and documents the change in the changelog.

Changes:

  • Add a new comma-splice classifier (comma_splice_watch / comma_splice_resumptive) plus a comma_splice_rate report, and upgrade prose scanning to be fence/directive-aware with paragraph-start/end tracking.
  • Extend ml_repair.py to apply the new splice repair using ml_metrics.split_resumptive_splices.
  • Add fixture-based unittest coverage and document the update in CHANGELOG.md.
File Description
experiments/​ml-benchmark/​scripts/​test_ml_lints.py New unittest fixtures covering round-4 splice repair and updated prose scanning/terminal punctuation behavior.
experiments/​ml-benchmark/​scripts/​ml_repair.py Adds a 4th deterministic repair (comma-splice split) and wires it to the lint classifier output.
experiments/​ml-benchmark/​scripts/​ml_metrics.py Adds splice detection + rate reporting; rewrites prose scanning to properly handle labels, directive bodies, and CommonMark-style fence closing.
CHANGELOG.md Documents the new Malayalam benchmark tooling behavior under [Unreleased].

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread experiments/ml-benchmark/scripts/ml_metrics.py
Comment thread experiments/ml-benchmark/scripts/ml_repair.py Outdated
…hole in a docstring

Addresses Copilot's review of #331. The terminal-punctuation lint accepted a
trailing comma before a code or math cell, although its own comment, rule 14
and the editor's ml#7 convention allow a comma only before a list it opens.
It now flags that case; the colon repair still leaves it to the lint, since
appending would write ",:". No such line occurs in any ml corpus (60+ files),
so no current output changes. A test covers both cases, and the docstring
reference lecture-python-programming.ml#23 no longer wraps mid-name.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mmcky
mmcky merged commit d9b243c into main Sep 28, 2026
1 check passed
@mmcky
mmcky deleted the ml/splice-lint-repair branch September 28, 2026 04:52
mmcky added a commit that referenced this pull request Sep 28, 2026
…anded (#332)

Refresh of the ml calibration bullet (round 4 on
lecture-python-programming.ml#23 applied and merged as 7e373de; #331's
splice lint + repair landed; next steps in order) and a Recently landed
entry for #331, plus the session log 2026-09-28-ml-round4. Records that
#331 is validated offline only: no regeneration test yet.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
mmcky added a commit that referenced this pull request Sep 28, 2026
…round-5 pandas seed arm (#333)

* ml: splice-repair regeneration arm — 18 fresh drafts, blind judges, precision screen

The regeneration test for #331: 18 fresh drafts at d9b243c (numpy,
functions, matplotlib x6), each repaired with ml_repair before (1e6936f)
and after (d9b243c) #331. The split is preferred 68 : 8 by a
reference-free Opus 5.5 judge and 60 : 21 by the style judge against the
editor's reviewed text; a two-reader screen passes all 81 split
paragraphs. The splice rate falls 17.3 -> 10.7 (numpy) and 18.3 -> 14.2
(functions) per 100 lines, still above the editor's 2.8. Raw drafts,
corpus, judgements, scores, screen verdicts and scripts archived.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ml: round-5 seed arm — three pandas drafts at v0.29.3, draft 1 sent as ml#26

Three drafts at 16e50f6 (@v0), chosen by lint on the residue no repair
touches; ml_repair at d9b243c applied and disclosed on the PR (8 -ഉം
commas, 1 splice split, 2 colons). Checks per draft in checks.jsonl.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ml arm: read and write the arm scripts as UTF-8; judge.mjs defaults to the model it ran with

Addresses Copilot's review of #333: every open() in make_items, sites,
summarise and repair_and_score now names encoding='utf-8' (Malayalam
text breaks on a non-UTF-8 platform default), and judge.mjs defaults to
claude-opus-5-5, the model the archived judgements came from
(JUDGE_MODEL still overrides). No archived output changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
mmcky added a commit that referenced this pull request Sep 28, 2026
…xception merged, Opus 5.5 scoped (#336)

* dev: STATE.md + log — #331 regeneration test, round 5 open, rule-19 exception merged, Opus 5.5 scoped

Refreshes the ml calibration bullet (the regeneration test in #333 is done;
only the rule-19 exception shipped in #335, with the three deletions
measured and rejected; round 5 is open as lecture-python-programming.ml#26)
and adds Recently landed entries for #333 and #335, plus a later-the-same-
day section in log 2026-09-28-ml-round4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* dev log: state the arm-D comparison with percentages and each arm's draft count

Addresses Copilot's review of #336: arm A had 36 drafts and arm D 24, so
the comparison now reads 42% (15 of 36, arm A) to 75% (18 of 24, arm D).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants