pgc_ledger.py merge prints one summary about the majors field, and that summary is structurally incapable of showing the only defect anyone has ever hit in it. The defect has now occurred twice in three hours, to the same person, with a written note about it in between.
The mechanism
test/pgc_ledger.py:473
majors = sorted(set().union(*(v[0] for v in rows.values())) if rows else set())
print(f" majors the ledger now claims rows for: {', '.join(majors) or 'none'}")
It is a UNION over rows. Merge eight rows carrying {18} into a ledger of 1209 rows carrying {15,16,17,18,19} and the union is unchanged, so the line is byte-identical on a correct merge and an incorrect one.
A union cannot represent per-row variation. It is the one statistic that cannot see a minority set, and it is the only one the merge emits.
Both occurrences
|
rows written |
against |
caught by |
| #1041 |
12 at 18 |
934 at 15;16;17;18;19 |
CI's suites (PG 17) leg |
| #1042 (in progress) |
8 at 18 |
1209 at 15;16;17;18;19 |
uniq -c over the whole column, locally |
Both were a single PG18 run merged, so the merge stamped the major it observed. In the first, every suite in that PG17 leg PASSED and the only failing condition was the majors field. In both, the merge printed majors ... 15, 16, 17, 18, 19.
The operator was not ignoring the output. The output agreed with them. A roll-up that cannot represent the failure is worse than no summary, because it actively confirms the wrong answer.
The distribution is the same information without the collapse
uniform 1209 15;16;17;18;19 ONE line
the defect 1209 15;16;17;18;19 + 8 18 TWO lines
One uniq -c separates them, and the tool already has the data — v[0] is the per-row major set it is unioning away.
Suggested shape
Print the distribution rather than the union, and say when it is not uniform. Same move the repo already makes elsewhere — --emitters printing the classification, the dynamic bucket printed rather than assumed — print what was classified, not the roll-up.
Whether merge should REFUSE a non-uniform result is a genuine design question and should not be settled by this issue — but the case against refusing is weaker than the obvious citation makes it look, and the difference is one word.
pgc_ledger.py:141 records "6367 of 6472 checks are identical on both majors and 105 exist on exactly one." That is a measurement of the CORPUS, not of the LEDGER. The ledger covers four suites (differential, harness_selftest, native_join_runtime_filter, native_join_vector_agg) and holds 0 rows with a non-full major set — all 1209 carry 15;16;17;18;19. So:
|
|
| checks in the corpus existing on exactly one major |
105 |
| such rows in the committed ledger today |
0 |
| rows refusing non-uniformity would forbid today |
0 |
A reader meeting "105" next to a refuse-or-report question will read it as 105 rows would be refused, and argue against refusing on that basis. The real number today is zero. The 105 is a statement about what could appear if the ledger's coverage grew to those suites, not about what is in it. Caught by @OffgridwithJD.
A distinguishing heuristic, and it is currently UNFALSIFIABLE rather than merely unverified. The guess: the defect writes rows whose major set equals the majors of the run just merged, whereas a legitimate divergence is a check missing on one major (4 of 5) rather than present on exactly the 1 that ran.
Since the ledger holds zero non-uniform rows, there are no instances on the current tree to distinguish it against in either direction. It is not contradicted by the data; there is no data. Anyone building on it must first produce a tree where a legitimate divergence exists — otherwise the heuristic will be validated against the defect alone, which every rule fits.
The GATE already refuses a check the ledger has never seen on the major being run — that is what caught #1041. So the merge printing a non-uniform distribution and the gate refusing it are the same fact observed at two moments, and the merge is simply the earlier and cheaper one. The gate catching it costs a full CI matrix; the merge catching it costs a line.
Why this is a tooling issue and not an operator one
Same person, same defect, three hours apart, with the lesson written down in between. Two occurrences under those conditions is the argument that memory is the wrong mechanism here, and it is stronger than either instance alone.
🤖 Generated with Claude Code
https://claude.ai/code/session_01NhwXKAgSmYDUjteWkfajHK
pgc_ledger.py mergeprints one summary about the majors field, and that summary is structurally incapable of showing the only defect anyone has ever hit in it. The defect has now occurred twice in three hours, to the same person, with a written note about it in between.The mechanism
test/pgc_ledger.py:473It is a UNION over rows. Merge eight rows carrying
{18}into a ledger of 1209 rows carrying{15,16,17,18,19}and the union is unchanged, so the line is byte-identical on a correct merge and an incorrect one.A union cannot represent per-row variation. It is the one statistic that cannot see a minority set, and it is the only one the merge emits.
Both occurrences
1815;16;17;18;19suites (PG 17)leg1815;16;17;18;19uniq -cover the whole column, locallyBoth were a single PG18 run merged, so the merge stamped the major it observed. In the first, every suite in that PG17 leg PASSED and the only failing condition was the majors field. In both, the merge printed
majors ... 15, 16, 17, 18, 19.The operator was not ignoring the output. The output agreed with them. A roll-up that cannot represent the failure is worse than no summary, because it actively confirms the wrong answer.
The distribution is the same information without the collapse
One
uniq -cseparates them, and the tool already has the data —v[0]is the per-row major set it is unioning away.Suggested shape
Print the distribution rather than the union, and say when it is not uniform. Same move the repo already makes elsewhere —
--emittersprinting the classification, thedynamicbucket printed rather than assumed — print what was classified, not the roll-up.Whether merge should REFUSE a non-uniform result is a genuine design question and should not be settled by this issue — but the case against refusing is weaker than the obvious citation makes it look, and the difference is one word.
pgc_ledger.py:141records "6367 of 6472 checks are identical on both majors and 105 exist on exactly one." That is a measurement of the CORPUS, not of the LEDGER. The ledger covers four suites (differential,harness_selftest,native_join_runtime_filter,native_join_vector_agg) and holds 0 rows with a non-full major set — all 1209 carry15;16;17;18;19. So:A reader meeting "105" next to a refuse-or-report question will read it as 105 rows would be refused, and argue against refusing on that basis. The real number today is zero. The 105 is a statement about what could appear if the ledger's coverage grew to those suites, not about what is in it. Caught by @OffgridwithJD.
A distinguishing heuristic, and it is currently UNFALSIFIABLE rather than merely unverified. The guess: the defect writes rows whose major set equals the majors of the run just merged, whereas a legitimate divergence is a check missing on one major (4 of 5) rather than present on exactly the 1 that ran.
Since the ledger holds zero non-uniform rows, there are no instances on the current tree to distinguish it against in either direction. It is not contradicted by the data; there is no data. Anyone building on it must first produce a tree where a legitimate divergence exists — otherwise the heuristic will be validated against the defect alone, which every rule fits.
The second half, raised by @OffgridwithJD
The GATE already refuses a check the ledger has never seen on the major being run — that is what caught #1041. So the merge printing a non-uniform distribution and the gate refusing it are the same fact observed at two moments, and the merge is simply the earlier and cheaper one. The gate catching it costs a full CI matrix; the merge catching it costs a line.
Why this is a tooling issue and not an operator one
Same person, same defect, three hours apart, with the lesson written down in between. Two occurrences under those conditions is the argument that memory is the wrong mechanism here, and it is stronger than either instance alone.
🤖 Generated with Claude Code
https://claude.ai/code/session_01NhwXKAgSmYDUjteWkfajHK