Skip to content

test(avro): pin the nominal-resolution doctrine — 391→431 mutants killed, 14 fewer examples - #114

Merged
kryptt merged 7 commits into
mainfrom
test/avro-mutant-kills
Sep 18, 2026
Merged

kryptt merged 7 commits into
mainfrom
test/avro-mutant-kills

Conversation

@kryptt

@kryptt kryptt commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Base: main (ad30b87). Rebased 2026-09-18 — no stacked-merge hazard remains, and the branch's merge-base was already a real ancestor of main, so this was an ordinary git rebase main, not an --onto recovery.

Four merges landed while this was open: #107 (cached nominal name index), #109 (kyo brace-for-stryker), #111 (cross-module mutant kills), #92 (zio ecosystem optics). One of them, #107, rewrote the exact code this PR tests — it deleted sameFieldName and the fuzzy nominalIndex scan and replaced them with normalisedName + a weak-keyed normalisedNameIndex cache, and made totalNominalIndex private[avro].

That rewrite cost this PR nothing, by construction. Every artifact here was written against the public seam AvroWalk.fieldNameAt, never against the internals, precisely so it would survive a rewrite of them. Not one line of test code needed changing: AvroNominalDoctrineSpec's doctrine oracle agrees with #107's new implementation cell for cell over the same ~66k-cell corpus. The only edit the rebase required was to the spec's covers: doc block, which named symbols #107 deleted and line numbers it moved (commit cb4edec) — a comment, not an assertion.

Two oracles now coexist on purpose, and both are wanted. #107's NominalResolutionOracle + vulcan/NominalResolutionParitySpec freeze the OLD ALGORITHM verbatim and ask "does the new rung agree with the one it replaced" — a refactor gate, which cannot fire if the algorithm was already resolving to the wrong field. This PR's oracle is written from the DOCTRINE without reading the implementation and asks "does the rung agree with what the rung is FOR". Parity cannot detect a wrong doctrine; that is the #104 failure mode. They collide in neither name nor fixture (main's is a top-level object, this one's is private to its spec), so nothing was renamed.

An improve-test-leverage run over the avro module: triage → 5 test artifacts → consolidation → law-scan. Only */src/test/ changed; nothing was promoted into laws/src/main (the law gate stops for you — the queued candidates are at the bottom).

The doctrine blind spots found and closed

#98 added an all-or-nothing nominal rung to AvroWalk: resolve .field(_.x) by NAME when every case field maps to a distinct schema field, else abstain to position. Four clauses. Mutation testing showed only one of them was pinned:

Clause Before After
1. Totality pinned (FpVisit / FpRow / Comp fixtures) pinned
2. Injectivity not pinned — no fixture made two case fields resolve to one slot; the seen scan had 2 survivors pinned
3. Exact-beats-normalised not pinned — the exact-match fast path was provably redundant in every fixture pinned
4. Ambiguity ⇒ abstain passed for the wrong reason — the sole fixture (ResolutionResidualSpec "at3"/Tw) also had an unmapped second field, so totality forced the abstention and masked ambiguity detection; 3 survivors pinned

This is not cosmetic: #104 recorded an experimental alias rung that silently re-aimed a case the rung gets right while all 212 avro tests stayed green. A mis-targeted optic satisfies get-put, put-get, put-put and modify fusion perfectly, which is why the original bug shipped.

AvroNominalDoctrineSpec now states the rung as a declarative oracle written from the doctrine (deliberately not a frozen copy of the implementation — parity against the code cannot detect a wrong doctrine, which is exactly #104's failure mode) and checks it exhaustively over ~66k cells. Threshold cells (a duplicate landing on slot 0, a fuzzy ambiguity whose first hit is at index 0 vs ≥1, declIdx exactly equal to the field count) are why it enumerates instead of generating.

Mutation numbers — and why they cannot be re-measured on this base

The figures below were measured on the pre-#107 tree (2e0b6b9 base, branch tips 035d70d / 8b1b184). They are reported as history, not as a claim about main today:

Killed Survived NoCoverage CompileError Ignored
baseline (2e0b6b9, PR #105 base) 391 78 17 6 122
after phases 1–5 (035d70d) 431 38 17 6 122
after consolidation (8b1b184) 431 38 17 6 122

80.45% → 88.68% total, 83.37% → 91.90% of covered, on that tree. Diffed by the full (file, line, column, mutator, replacement) key: 40 Survived→Killed, ZERO Killed→anything, zero Survived→Timeout flips. Every flipped key was claimed by exactly one artifact — predicted == measured, 40/40. Consolidation was diffed the same way: ONLY_IN_A 0 · ONLY_IN_B 0 · REGRESSIONS 0 · GAINS 0, measured twice.

⚠️ avroIntegration no longer mutates at all — on main, not here

Re-running sbt -batch 'set ThisBuild/tlFatalWarnings := false' 'project avroIntegration' stryker after the rebase fails before producing a report. It fails identically on origin/main with this branch nowhere in sight, so it is a main regression, not something the rebase introduced. Two independent blockers, in order:

  1. AvroWalk.scala, totalNominalIndex (added by perf(avro): one cached name index per schema, not a scan per case field (#103) #107) —
    val here = if index != null then index else normalisedNameIndex(record).
    Under -Yexplicit-nulls the local's type is inferred from the != null flow test. Mutate the condition and the type is JMap[String, Integer] | Null, so here.get(...) no longer compiles, and stryker4s reports UnableToFixCompilerErrorsException: No mutants were removed in AvroWalk.scala even though there were 1 compile errors. Verified fix (not included here — this PR is test-only): ascribe the local and use .nn, i.e. val here: JMap[String, Integer] = if index != null then index.nn else normalisedNameIndex(record). Same shape as the fix(kyo): brace the valuePrism extension so stryker4s can mutate the module #109 fix in kyo. Applied locally, this blocker disappears.
  2. AvroBinaryCursor.scala — with (1) fixed, the run then dies in stryker4s's rollback with org.scalameta.invariants.InvariantFailedException: invariant failed (cases should be non-empty), rolling back 37 compile-error mutants. Also reproduces on plain main + fix (1). Not diagnosed further here.

So the honest statement of this PR's mutation effect against current main is: unmeasured, because the measurement is broken upstream of this branch. The two blockers deserve their own PR (the #109 pattern); I have deliberately not smuggled a production change into a test-only PR to get a number. What is measured on the current base is below — the suite runs, passes, and shrank.

Kills per line

Artifact Kills Body lines lines/kill
AvroNominalDoctrineSpec — exhaustive doctrine oracle 12 68 5.7
AvroJsoniterSpec — parameterised malformed-JSON negative fixtures 7 13 1.9
AvroWalkSpec — decoded-map keys + union diagnostic payloads 6 46 7.7
AvroBytesSpec + Twin — byte-cursor refusal identities, strictness policy 11 83 7.5
AvroWriteCorrectnessSpec — array framing survives a write 4 32 8.0
total 40 242 6.05

The ≤8 lines/kill gate bit on three of the five as submitted (11.7 / 8.9 / 9.0). Rather than revert 21 verified kills, each was trimmed semantics-preservingly (shared at/field/branch/tagAt helpers; hand-rolled match + specs2.execute.Failure blocks replaced by direct === Left(...) equalities, several of which assert strictly MORE than what they replaced) and stryker re-run: outcome-identical, same 40 keys, 0 regressions. No assertion was weakened to shrink a denominator.

Every row asserts which AvroFailure comes back, not isLeft — each removed guard degrades to BinaryParseFailed, a Left either way. The write-side rows assert bytes on purpose: every write decodes to the right value under every mutation of the three framing guards, so the Optional-law property block at the foot of the same file cannot see any of it.

suite_test_count delta — the suite SHRANK

revision >> registrations test LOC
2e0b6b9 baseline (pre-#107) 173 4,952
035d70d after phases 1–5 177 5,201
8b1b184 after consolidation 159 5,122

Re-measured on the current base, as specs2 example counts for the whole avroIntegration module (which now also carries #107's three new specs): origin/main 228, this branch 214 — −14, exactly the delta above, so the rebase preserved it.

−14 examples for +40 kills (kills measured on the pre-#107 tree; see the section above). Consolidation alone: −18 examples, −79 lines, zero kill regression — safe before it was measured, by attribution: stryker records the killing test in each mutant's statusReason, and of avro's 210 registered tests only 65 are ever the credited killer of anything. Every folded example was in the uncredited 145, so each killed mutant's credited killer survived the fold untouched. Seven folds, 27 examples → 9; three of them strengthened an assertion (e.g. the render half of the enum/bytes/fixed fold now asserts the whole document, avroToJson(record) === json, instead of three field lookups).

LOC is +170 vs baseline while the example count is below it: phases 1–5 added 242 lines of dense tests and consolidation removed 79 — 4.25 lines per additional kill with 14 fewer examples.

Classified equivalent (with checkable reasons)

  • ~6 in AvroWalk — >=/== boundary mutants on cursors incremented by exactly 1: the index can never skip the bound, so both operators are observationally identical.
  • AvroPrismMacro.scala, all 17 NoCoverage — quoted-macro code that expands at compile time. Stryker mutates runtime bytecode, so it structurally cannot reach it; same blind spot that scores generics at 0%. Noted, not chased.
  • AvroFocus's 4 survivors — verified to be fast-path mutants whose slow path computes the same value.
  • AvroBinaryCursor's end-of-stream guards were flagged "investigate, not resolved" in triage; 11 of them are now killed by the byte-cursor artifact, the remainder stay open.

Remaining: 38 Survived (AvroBinaryCursor 18, AvroWalk 14, AvroFocus 4, AvroCodec 1, AvroRecordTraversal 1) and 17 NoCoverage (all AvroPrismMacro).

Queued law candidates — for your decision, nothing promoted

Phase 7 scanned ~683 example registrations and 143 forAll shapes across 91 spec files in 10 modules, plus all 59 files of laws/src/main, for properties that are universal for an abstraction rather than specific to avro. Two candidates survive vetting; both were measured with scratch probes (kept out of the repo, in the session scratchpad).

Q1 — TargetingLaws / TargetingTests (issue #99)

The extrinsic law. For an optic claimed to target a named location n, given a reference accessor the carrier itself supplies (never derived from the optic):

  • T1 optic.getOption(s) == Some(ref(s, n))
  • T2 ref(optic.replace(b)(s), n) == b
  • T3 ∀ m ≠ n: ref(optic.replace(b)(s), m) == ref(s, m) — no collateral write

Measured, not argued (ScratchTargetingProbeSpec, 5/5 green, avro Three fixture — schema {a, computed, b, c} vs case class {a, b, c}): two optics onto the same source, .field(_.b) and .fieldNamed("computed") (where #95's positional resolution used to send _.b). Both satisfy all five intrinsic equations — modify-identity, compose-modify, replace-overwrite, get-put, put-get — which is the blindness claim, now a measurement rather than a story. The correct one passes T1/T2/T3; the mis-targeted one is rejected by all three. T3 has independent teeth: an optic that hits the right slot and clobbers a sibling passes T1 and T2 and fails T3 only.

Universality. The hazard is carrier-wide, not an avro quirk: circe and jsoniter resolve by the literal Scala name (a JSON key transform mis-targets identically), zio.schema's EoAccessorBuilder and kyo's Record.lens[F]("name") are the same shape. Every one of those carriers already ships a native by-name accessor to serve as ref.

The hard part, honestly. The law needs an independent notion of "the right target", which the optic cannot supply. Three sourcings, and this is the decision:

  1. Fixture-declared (what the probe does): the fixture states the intended name→slot mapping. Sound, zero new machinery, but the evidence quantifies over hand-written fixtures, so it catches regressions rather than unknown shapes.
  2. Carrier-native accessor (GenericRecord.get(name), hcursor.downField, Record's key lookup): independent only if the Scala-name→carrier-name mapping comes from the fixture. Derive that mapping from the same name-matching machinery under test and the law is vacuous — this is the one way to get this wrong.
  3. Derived, via avro: differential codec probe — ask the codec which slot a field lands in, instead of guessing from names #100's differential codec probe: encode two sentinels, diff the slots, observe where a case field actually lands. That is ground truth, needs no declaration, and scored 7 true positives / 0 false positives / 1 false negative on the full DivergentCodecs + ResolutionResidualSpec matrix. avro: differential codec probe — ask the codec which slot a field lands in, instead of guessing from names #100 is the oracle laws: TargetingLaws / TargetingTests — the only thing that can catch a resolution mechanism's own mistake #99 wants — it upgrades the reference from hand-supplied to derived, and it is exactly what makes the ruleset registrable at shapes nobody wrote a fixture for.

There is precedent inside laws/ already: GetterLaws.reference and UnfoldLaws.reference are extrinsic-reference laws, fixture-supplied, and registered today. Q1 is that pattern generalised, not a new species.

Cost: ~45 lines in laws/src/main/scala/.../laws/eo/TargetingLaws.scala + a TargetingTests ruleset; per-carrier registration needs a reference accessor and a sibling list, so this is not a zero-test-file-change promotion. Buys: the only mechanism that can catch a resolution mechanism's own mistake, in five carriers, for every future rung (#104's alias rung would have been caught by T1 with a crossed-alias fixture). MiMa vacuous (tlMimaPreviousVersions := Set.empty).

Q2 — ForgetfulFoldLaws — fold/traverse coherence

ForgetfulFold[F] is the typeclass every optic read goes through — 6 instances (Tuple2, Either, Affine, Direct, Forget[F], MultiFocus[F]) — and it has no law trait at all. The two shipped fold laws are stated at the optic level (FoldLaws, FoldMapHomomorphismLaws) and are both satisfied vacuously by a fold that returns Monoid.empty for everything.

FF.foldMap(f, fa) == FT.traverse[Const[M, *]](fa, a => Const(f(a))).getConst

stated at the free monoid List[A]. This is the classical Foldable/Traverse coherence, and it doubles as the naturality witness that ForgetfulTraverseLaws's own scaladoc flags as absent ("requires an applicative-transformation fixture, which in turn needs a second concrete Applicative" — Const[M, *] is that second applicative, and cats already ships it).

Measured (ScratchLawScanSpec, 14/14 green): holds on all 6 carrier shapes reachable from a registered fixture (Tuple2/Lens, Either/Prism, Affine/Optional, Affine/AffineFold, MultiFocus[PSVec]/Traversal, Forget[List]/Fold), 100 samples each, zero counterexamples. Minimally unlawful instance: the const-empty ForgetfulFold[Affine] — it satisfies the monoid-homomorphism law and the foldMap(const(mempty)) == mempty law (both vacuously, measured over 200 cases each), so no registered ruleset today rejects it; this law rejects it. Not stateable for Direct, which has no ForgetfulTraverse[Direct, Applicative].

Cost: ~20 lines of law + ~20 of ruleset, and one edit to CheckAllHelpers.checkAllForgetfulTraverseFor — its 9 existing call sites then inherit it with zero test-file changes (the stage-3 payoff). All 9 already have the ForgetfulFold evidence in scope. MiMa vacuous. Buys: any mutation of a ForgetfulFold instance body that preserves the homomorphism shape — returning empty, folding the leftover instead of the focus, dropping an element of a multi-focus — currently survives every registered ruleset. Predicted kills unmeasured: needs a core stryker run, which this run (scoped to avro) did not do.

Q3 — optic-level read-after-write coherence: scanned and rejected

foldMap(List(_)) ∘ modify(f) == map(f) ∘ foldMap(List(_)) looked like the missing bridge — TraversalLaws has no read-side law at all and FoldMapHomomorphismLaws has no write side, so nothing today relates the two. It is not worth promoting, on two measured grounds: it holds vacuously for the const-empty fold that Q2 rejects (Nil == Nil.map(f)), so it is strictly weaker than Q2; and its only witness — a prism whose reverseGet lands outside its own match — is already rejected by PrismLaws.roundTripOtherWay. Both facts are pinned in the probe. It also needs a domain condition the shipped fixtures already document (the registered evenPrism site restricts Int to ±10000 because doubling overflows), which is a sign it is a property of the fixture, not of the abstraction.

Non-law finding: a registration gap

jsoniter, zio and kyo register zero checkAll sites between them (tests 104, circe 5, avro 5, schemes-laws 1). Their optics — jsoniter's cursor faces, EoAccessorBuilder, Record.lens, the StructureValues kits — are exactly the shapes the shipped rulesets were written for. That is stage-3 coverage available with no new laws at all, and it is where the next run's cheapest leverage sits, alongside the 127 still-uncredited avro examples (clustered in ConfluentReaderSpec 12, AvroWriteCorrectnessSpec 12, circe/AvroJsonSpec 8, jsoniter/AvroJsoniterSpec 11).

Gates

All on JDK 25 (Temurin 25.0.4.1), after the rebase onto ad30b87:

  • sbt "++ 3" compile test — root aggregate, all 10 test modules green, 0 failures / 0 errors. kyoIntegration confirmed present in root/aggregate and exercised (KyoOpticsSpec, RecordOpticsSpec, SchemaOpticsSpec, StructureOpticsSpec, StructureValuesLawsSpec all ran) — so the JDK-21 silent-drop trap did not fire.
  • avroIntegration — Passed 214, Failed 0, Errors 0 (212 passed, 2 skipped: NominalResolutionCostSpec's wall-clock gate, disarmed unless -Deo.costGate=true). AvroNominalDoctrineSpec, NominalResolutionParitySpec, NominalNameIndexCacheSpec and NominalResolutionCostSpec all green together.
  • sbt benchmarks/compile — clean (outside the root aggregate, breaks silently).
  • sbt mimaReportBinaryIssues — clean.
  • sbt "docs/mdoc; docs/laikaSite" — clean.
  • Format gate run last, with the stale-cache precaution (target/streams/**/scalafmt*, **/scalafix* deleted first): scalafmtCheckAll re-read 21 source groups and scalafixAll --check re-ran over 39 sources — both clean, working tree clean.
  • project avroIntegration; stryker — fails on main itself; see the ⚠️ section above.

What the rebase changed

  • One of the branch's seven commits, 0b6a3e9 docs(qa): drop the stale "not scored" caveats for avro and jsoniter, became empty and was dropped: test: 29 mutants killed in −2 net test lines — jsoniter oracle rewrite, circe bounds property, core exists/Index, and the equivalent-mutant catalogue #111 had already retired both caveats on main, and did it better — it replaced them with accurate notes ("Scores fine (~2 min)…", "Mutates clean end to end…") rather than blanking the cells. Both conflicting hunks in site/docs/quality-assurance.md and site/tools/gen-qa-report.py were resolved in main's favour, so nothing this commit intended is missing. (Note that main's "Scores fine (~2 min)" note is itself now stale, for the reason in the ⚠️ section — a main problem, not one to fix from a test branch.)
  • The other six commits survived verbatim; cb4edec was added for the covers: doc block.
  • git diff origin/main..HEAD is 10 files, all under avro/src/test/ — no production code, no docs, no build files.

🤖 Generated with Claude Code

https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

🚀 Cloudflare Pages preview for test/avro-mutant-kills is live:

https://3bc924a5.cats-eo-docs.pages.dev

Branch alias: https://test-avro-mutant-kills.cats-eo-docs.pages.dev

Built from commit cb4edecf82af6af9cf420fe4b0f3097c5b511635 · updated on every push.

@github-actions

github-actions Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Benchmark A/B

Allocation (B/op) — authoritative

Benchmark params base head Δ
AvroVulcanBench.encode_vulcanRaw - 1,208.0 1,240.0 +2.6%
AvroVulcanBench.encode_bridged - 1,240.0 1,272.0 +2.6%
AvroJsonBridgeBench.eoWideToAvro - 6,584.0 6,744.0 +2.4%
AvroBytesBench.eoGraftPayload - 720.0 704.0 -2.2%
AvroJsonBridgeBench.eoClickToJson - 4,000.0 3,952.0 -1.2%
AvroJsonBridgeBench.eoClickToAvro - 9,432.0 9,400.0 -0.3%
OrderAvroBench.monocleModifyStreet size=64 20,904.2 20,880.2 -0.1%
OrderAvroBench.eoModifyNames size=512 97,404.4 97,430.5 +0.0%
OrderAvroBench.naiveReadStreet size=512 69,801.7 69,802.0 +0.0%
OrderAvroBench.monocleModifyStreet size=512 169,065.9 169,065.5 -0.0%
OrderAvroBench.monocleModifyNames size=64 39,867.3 39,867.4 +0.0%
OrderAvroBench.naiveModifyStreet size=512 169,057.7 169,057.4 -0.0%
OrderAvroBench.eoModifyNames size=64 12,584.1 12,584.2 +0.0%
OrderAvroBench.eoModifyNames size=8 2,160.0 2,160.0 -0.0%
OrderAvroBench.eoModifyStreet size=512 328.0 328.0 -0.0%
OrderAvroBench.eoReadStreet size=512 88.0 88.0 +0.0%
OrderAvroBench.monocleReadStreet size=512 69,802.0 69,802.0 +0.0%
AvroBytesBench.prunedReadPartner - 1,592.0 1,592.0 +0.0%
AvroBytesBench.eoReadPartner - 480.0 480.0 +0.0%
OrderAvroBench.monocleReadStreet size=64 8,848.2 8,848.2 -0.0%
38 more benchmarks
Benchmark params base head Δ
AvroJsonBridgeBench.naiveClickToJson - 4,696.0 4,696.0 +0.0%
AvroBytesBench.eoReadCountry - 520.0 520.0 -0.0%
AvroBytesBench.eoSliceGraftPayload - 1,192.0 1,192.0 +0.0%
AvroJsonBridgeBench.naiveWideToJson - 4,376.0 4,376.0 -0.0%
OrderAvroBench.eoModifyStreet size=64 328.0 328.0 +0.0%
OrderAvroBench.naiveModifyNames size=64 27,936.3 27,936.3 -0.0%
AvroJsonBridgeBench.eoWideToJson - 1,472.0 1,472.0 +0.0%
OrderAvroBench.monocleModifyNames size=512 382,759.4 382,759.5 +0.0%
OrderAvroBench.eoModifyStreet size=8 328.0 328.0 +0.0%
OrderAvroBench.monocleModifyStreet size=8 2,992.0 2,992.0 +0.0%
OrderAvroBench.naiveModifyStreet size=8 2,968.0 2,968.0 -0.0%
AvroBytesBench.naiveModifyPartner - 7,536.0 7,536.0 +0.0%
AvroVulcanBench.rootGet_native - 600.0 600.0 +0.0%
AvroBytesBench.eoModifyCountry - 3,200.0 3,200.0 +0.0%
OrderAvroBench.naiveModifyStreet size=64 20,880.2 20,880.2 -0.0%
AvroVulcanBench.fieldGet_native - 432.0 432.0 -0.0%
AvroBytesBench.naivePassthroughPayload - 10,600.1 10,600.1 +0.0%
OrderAvroBench.monocleReadStreet size=8 1,208.0 1,208.0 -0.0%
OrderAvroBench.naiveReadStreet size=8 1,208.0 1,208.0 +0.0%
AvroVulcanBench.rootGet_bridged - 1,472.0 1,472.0 +0.0%
AvroBytesBench.naiveModifyCountry - 7,616.0 7,616.0 +0.0%
AvroBytesBench.prunedReadCountry - 1,976.0 1,976.0 +0.0%
AvroVulcanBench.decode_bridged - 880.0 880.0 +0.0%
AvroBytesBench.naiveReadPartner - 4,264.0 4,264.0 +0.0%
AvroVulcanBench.decode_vulcanRaw - 880.0 880.0 +0.0%
AvroBytesBench.naiveReadCountry - 4,256.0 4,256.0 +0.0%
OrderAvroBench.naiveModifyNames size=8 3,752.0 3,752.0 -0.0%
OrderAvroBench.eoReadStreet size=8 88.0 88.0 -0.0%
AvroJsonBridgeBench.naiveClickToAvro - 3,928.0 3,928.0 -0.0%
AvroJsonBridgeBench.naiveWideToAvro - 3,504.0 3,504.0 +0.0%
OrderAvroBench.eoReadStreet size=64 88.0 88.0 +0.0%
OrderAvroBench.naiveModifyNames size=512 226,283.5 226,283.5 +0.0%
AvroVulcanBench.fieldGet_bridged - 432.0 432.0 -0.0%
OrderAvroBench.naiveReadStreet size=64 8,848.2 8,848.2 +0.0%
OrderAvroBench.monocleModifyNames size=8 5,400.0 5,400.0 +0.0%
AvroBytesBench.eoModifyPartner - 3,360.0 3,360.0 +0.0%
AvroVulcanBench.decode_native - 48.0 48.0 -0.0%
AvroVulcanBench.encode_native - 56.0 56.0 +0.0%
Timing (ns/op) — directional only, same-VM but shared runner
Benchmark params base head Δ
OrderAvroBench.monocleModifyStreet size=64 7,307.8 6,300.6 -13.8%
OrderAvroBench.eoModifyNames size=64 4,153.0 4,596.2 +10.7%
AvroJsonBridgeBench.eoWideToAvro - 1,004.2 1,069.8 +6.5%
AvroJsonBridgeBench.naiveClickToJson - 2,631.8 2,756.9 +4.8%
AvroVulcanBench.encode_vulcanRaw - 245.8 257.3 +4.7%
AvroBytesBench.prunedReadPartner - 549.7 574.6 +4.5%
AvroJsonBridgeBench.naiveWideToJson - 1,895.3 1,822.2 -3.9%
AvroBytesBench.eoReadPartner - 202.0 208.8 +3.4%
AvroJsonBridgeBench.eoWideToJson - 664.0 684.6 +3.1%
AvroBytesBench.eoSliceGraftPayload - 325.2 334.6 +2.9%
AvroBytesBench.eoReadCountry - 176.0 171.4 -2.6%
OrderAvroBench.eoModifyNames size=512 35,150.1 34,242.6 -2.6%
AvroVulcanBench.rootGet_native - 168.0 172.2 +2.4%
AvroVulcanBench.fieldGet_native - 99.0 96.6 -2.4%
AvroJsonBridgeBench.eoClickToJson - 2,885.4 2,815.6 -2.4%
AvroJsonBridgeBench.eoClickToAvro - 3,174.9 3,111.3 -2.0%
OrderAvroBench.eoModifyStreet size=8 120.7 122.6 +1.6%
OrderAvroBench.monocleModifyStreet size=8 982.3 997.4 +1.5%
OrderAvroBench.naiveModifyStreet size=8 984.2 970.1 -1.4%
OrderAvroBench.monocleModifyStreet size=512 56,347.7 55,545.3 -1.4%
AvroVulcanBench.rootGet_bridged - 439.7 445.7 +1.4%
AvroVulcanBench.decode_bridged - 223.6 226.5 +1.3%
OrderAvroBench.naiveReadStreet size=512 34,119.5 34,494.8 +1.1%
AvroVulcanBench.decode_vulcanRaw - 213.7 216.0 +1.1%
OrderAvroBench.eoModifyStreet size=512 122.3 121.2 -0.9%
AvroVulcanBench.encode_bridged - 263.3 265.6 +0.9%
OrderAvroBench.naiveModifyNames size=64 8,924.7 8,846.4 -0.9%
AvroBytesBench.eoGraftPayload - 159.4 160.7 +0.9%
OrderAvroBench.monocleReadStreet size=8 516.2 511.7 -0.9%
AvroBytesBench.naiveModifyPartner - 2,636.8 2,658.2 +0.8%
OrderAvroBench.monocleReadStreet size=64 4,423.3 4,388.7 -0.8%
OrderAvroBench.eoModifyStreet size=64 122.0 122.8 +0.7%
OrderAvroBench.eoReadStreet size=512 38.3 38.6 +0.6%
OrderAvroBench.naiveReadStreet size=8 518.1 521.2 +0.6%
OrderAvroBench.eoModifyNames size=8 597.4 600.7 +0.6%
AvroBytesBench.naivePassthroughPayload - 3,943.5 3,964.2 +0.5%
OrderAvroBench.naiveModifyNames size=512 69,418.7 69,066.0 -0.5%
AvroBytesBench.eoModifyPartner - 527.8 530.5 +0.5%
OrderAvroBench.naiveModifyStreet size=512 55,918.3 56,192.5 +0.5%
AvroBytesBench.naiveModifyCountry - 2,594.4 2,606.4 +0.5%
OrderAvroBench.naiveModifyStreet size=64 7,273.7 7,247.0 -0.4%
AvroBytesBench.prunedReadCountry - 740.1 742.6 +0.3%
OrderAvroBench.naiveModifyNames size=8 1,162.1 1,158.1 -0.3%
AvroBytesBench.naiveReadPartner - 1,760.2 1,765.8 +0.3%
AvroBytesBench.eoModifyCountry - 390.6 389.5 -0.3%
AvroBytesBench.naiveReadCountry - 1,665.2 1,669.4 +0.2%
AvroJsonBridgeBench.naiveWideToAvro - 1,003.0 1,005.2 +0.2%
AvroJsonBridgeBench.naiveClickToAvro - 1,491.3 1,488.3 -0.2%
AvroVulcanBench.fieldGet_bridged - 96.9 96.7 -0.2%
OrderAvroBench.monocleModifyNames size=512 101,432.8 101,594.2 +0.2%
OrderAvroBench.monocleReadStreet size=512 34,524.9 34,578.7 +0.2%
OrderAvroBench.eoReadStreet size=8 38.2 38.2 -0.1%
AvroVulcanBench.encode_native - 12.0 12.0 -0.1%
OrderAvroBench.monocleModifyNames size=64 10,362.1 10,368.0 +0.1%
OrderAvroBench.naiveReadStreet size=64 4,403.9 4,401.7 -0.1%
OrderAvroBench.monocleModifyNames size=8 1,472.7 1,471.9 -0.0%
OrderAvroBench.eoReadStreet size=64 38.2 38.2 +0.0%
AvroVulcanBench.decode_native - 18.6 18.6 -0.0%

base_sha: ad30b87ce2171fdc815fc72e8cb6e4eba6c988d7 · head_sha: cb4edecf82af6af9cf420fe4b0f3097c5b511635 · jdk: temurin-21 · runner: ubuntu-22.04 · jmh_params: -i 3 -wi 2 -f 1 -t 1 -foe true -prof gc -rf json · profile: pr:-i3-wi2-f1-t1-gc

kryptt and others added 7 commits September 18, 2026 18:13
…68 lines

PR #98 added an all-or-nothing nominal rung to `AvroWalk.fieldNameAt`: resolve
`.field(_.x)` by NAME when every case field maps to a distinct schema field,
otherwise abstain to position. Four clauses; only TOTALITY was pinned.
Injectivity, exact-beats-normalised, and ambiguity-implies-abstain had no
fixture that could tell a correct implementation from a broken one — the sole
ambiguity fixture (`ResolutionResidualSpec` "at3") passes for the wrong reason,
because an unrelated unmapped field forces totality-driven abstention first and
masks the ambiguity branch entirely.

That is not cosmetic. Issue #104 recorded an experimental alias rung silently
breaking a case the rung gets right while all 212 avro tests stayed green: a
MIS-TARGETED optic satisfies get-put, put-get, put-put and modify fusion
perfectly, so no optic law family can see this.

`AvroNominalDoctrineSpec` states the doctrine as a declarative oracle and checks
`fieldNameAt` against it EXHAUSTIVELY over ~66k cells: every injective field
list of length 1..3 from a 6-name pool (each name both a legal Avro field name
and a legal Scala identifier), crossed with every case-name list of length 0..3
WITH repetition, crossed with declIdx from -1 to arity+1. The oracle is written
from the doctrine, not copied from the implementation — parity against a frozen
copy of the code cannot detect a wrong doctrine, which is precisely the #104
failure mode.

Enumeration, not `Gen`: the discriminating cells are threshold cells a default
generator would essentially never produce — a duplicate resolution landing on
slot 0, a fuzzy ambiguity whose FIRST hit is at index 0 versus at index >= 1,
and `declIdx` exactly equal to the schema's field count.

Verified by applying each mutant to `AvroWalk.scala` and confirming the property
goes red: 291 (exact fast path), 301 / 304 (the injectivity scan), 169 / 171 /
173 (ambiguity => abstain), 213 (`sameFieldName`'s char inequality), 310
(`declIdx >= fields.size` at the exact boundary), and 242 / 244 / 245 / 249 (the
four disjuncts of the all-or-nothing precondition).

Deletes the subsumed `"fieldNameAt: non-record parent and out-of-range declIdx
both throw loudly"` example: it probed `declIdx = 99` against a 2-field schema,
a cell where `>=` and a `>` mutant agree, so it pinned neither boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
…in 13 lines

`AvroJsoniter.readRecord`'s strict exact-cover parse has three independent
refusals — unknown key, duplicate key, and the trailing arity check — plus
end-or-comma refusals in the array, map and byte-array readers. The module had
no fixture for any of the delimiter faults, and its one duplicate-key arm was
MASKED: appending `"i":42` makes 14 keys against 13 schema fields, so the arity
check refuses the document first and the `seen` scan is never exercised. Same
masking shape as the ambiguity fixture this branch's first commit replaced.

Every new arm is deliberately COUNT-PRESERVING:

  - the duplicate REPLACES a sibling key (`"l":...` becomes a second `"i":42`),
    leaving 13 keys with `l` unset, so `seen` is the only thing that can refuse;
  - the four delimiter faults close a nested record / array / map / bytes array
    with the wrong bracket, which the outer reader then resumes past, so the
    document parses end-to-end unless the inner reader refuses.

Kills 54 (`seen(field.pos) = true`), 57 and 59 (the `(field eq null) ||
seen(...)` guard and its operator), 71 (the record's objectEndOrCommaError), and
39 / 55 / 62 (the array / map / byte-array end-or-comma refusals).

Drops the now-dominated 14-key duplicate arm. The triage's proposed "unknown
JSON key" arm is NOT added: under both surviving guard mutants an unknown key
dereferences a null `field`, the NPE is caught by `parseSlice`'s `NonFatal`
handler, and the result is byte-identical to the real refusal — an equivalent
path, not a gap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
…mutants in 46 lines

Three gaps in the in-memory walker, all of the same species: the existing block
asserts the SHAPE of the answer and never its content.

Map keys. The map-walk example used a hand-built `LinkedHashMap[String, _]`, so
`asMap.get(name: String)` hits directly and the `direct == null` Utf8 retry is
dead code under test. A decoded map is keyed by `Utf8` — the shape production
code actually sees — and only the retry finds anything in it. The new arm puts
the record through the binary codec and probes every key of a 6-name pool
against map sizes 0 / 1 / 2 / 5, so present and absent probes both occur at
every size. Kills 322 (the Utf8 retry) and 338 (the PathMissing refusal, whose
removal returns `Right(null)`).

Null union branch. No fixture walked `UnionBranch("null")` into a null payload,
the one branch name for which a null focus is correct rather than a failure.
Kills 191.

Union diagnostics. `UnionResolutionFailed` carries the declared alternative
list, and every existing assertion matched it with a wildcard. Recovering that
list walks back to the nearest Field step, so it needs a union two records deep
to separate the parents cursor's `pIdx >= 0` guard from a `pIdx == 0` one (kills
183), and a path whose recovered Field step names an ARRAY rather than a union
to pin the empty-list degradation — asking a non-union for its alternatives
throws an AvroRuntimeException straight out of the walk (kills 254 and 256).

Folded into the existing per-shape composite blocks per this file's stated
consolidation convention: no new registrations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
…— kills 11 mutants in 83 lines

Every guard in `AvroBinaryCursor`'s path walk, when removed, lets the cursor run
on into the Avro runtime, which throws, which `locateFrom` catches and reports
as `BinaryParseFailed`. A `Left` either way — so the module's `isLeft`
assertions pin none of them, and the structured diagnostic the `.record` face
publishes degrades silently to "parse failed". These rows assert WHICH
`AvroFailure` comes back, never that one does.

Called at the `AvroBinaryCursor` seam (`private[avro]`, same package) rather
than through a prism: several rows need a path the drilling macros refuse to
build, and the strictness flag is not reachable from the public surface at all.

Kills 592 (Field step against a non-record), 500 and 597 (UnionBranch step
against a non-union, and the diagnostic's own non-union guard), 601 and 605
(the ordinal scan running off the end of the alternative list), 534 (the
`requested < 0` refusal under the LENIENT policy — the strict path refuses for a
different reason and masks it), 575 (locateElements' non-array terminal) and 506
(locateElements resolving its prefix strictly).

Two rows carry their own fixtures:

  - 542, the terminal-vs-interior union split. An interior union step must not
    anchor the returned span. Every downstream observable still round-trips
    under the mutant, because the mis-anchored span still opens on a decodable
    value; only `span.valueSchema` gives it away.
  - 440, `AvroPrism`'s byte-face read resolving its terminal union strictly. The
    existing union fixtures all mismatch INTO a decode error, which `=== None`
    cannot tell apart from a correct refusal. New `Twin` fixture: two branches
    with identical field shapes and different names, so `TwinA(100)` and
    `TwinB(100)` encode to the same bytes and a lenient read hands back a
    well-formed value of the wrong type. The lenient policy is correct for the
    graft/write path only.

The triage listed 440's witness as a null-branch payload with a trailing long
field. That does not discriminate: `decodeSlice` reads a LENGTH-BOUNDED slice,
and a null branch skips zero bytes, so the mutant hits EOF and Misses exactly
like the real path. Replaced with the `Twin` fixture above.

Also pins AvroTraversal 359 — `.each`'s prefix must resolve to an array or throw
loudly. The public macro rejects a non-array focus at compile time, so the guard
is only reachable at the internal seam, which is what a future drilling entry
point would use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
The traversal's write path has three guards nobody could see: the no-op guard
(`active.isEmpty || arity mismatch`), the canonical-splice vs re-frame split,
and the initial canonical flag the block walk threads. Mutating any of them
still produces a payload that DECODES to the right value, which is why the
Optional-law property block at the foot of this file — and every other
write-side assertion in the module — leaves all four alive. What changes is the
FRAMING: a payload that arrives blocked and leaves canonical, or arrives
multi-block and leaves single-block, is a rewrite of bytes the caller never
asked us to touch.

Two cells, both asserted at byte level, and neither is reachable from the
existing fixtures:

  - all-miss under NON-canonical framing (every element on the other union
    branch). The write must be the identity on the bytes. On canonical framing
    the mutants re-frame to something byte-identical, so only the blocked
    fixture separates them. Kills 364 and 361.
  - canonical MULTI-BLOCK framing plus a length-preserving edit. Spec-legal, and
    the only shape where an in-place splice and a whole-region re-frame differ:
    collapsing two positive-count blocks into one costs a count byte, so the
    payload shrinks while decoding identically. Every other fixture here is
    single-block, where a re-frame reproduces the input exactly. Kills 356 and
    586. Hand-assembled, since neither avro encoder emits it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
The mandatory consolidation phase. Seven folds, chosen by MEASURED redundancy
rather than by reading: stryker4s records the killing test in `statusReason`,
so the report itself says which examples carry kills. Of avro's 210 registered
tests only 65 are ever the credited killer of a mutant; every example folded
below is in the other 145, which makes zero kill regression a property of the
report and not a hope — the credited killer of each of the 431 killed mutants
survives every fold untouched. It was then measured anyway.

  1. AvroCodecDecoderReuseSpec — SEVEN properties differing only in the
     (generator, writer schema, reader schema) triple and in which decode entry
     point they called, collapsed to ONE. The triple is now data
     (`reuseScenarios`), and each scenario is put through BOTH entry points —
     the thread-local cache and the `threadLocalStorage = false` opt-out —
     where before each was checked by one example each. Strictly more coverage
     per shape.
  2. AvroFieldNamingSpec — nine examples to three. The four byte-face examples
     were one fixture on one code path (fixture guard, `.field` read, `.field`
     modify, `selectDynamic`); record face and nested descent were a second
     pair. The two `.fieldNamed` examples are duplicates of
     `AvroNominalResolutionSpec`'s, which asserts strictly more of the refusal
     message, so they go. `.fields` stays its own example: it is the one shape
     in the module where the FOCUS codec's names diverge from the parent's.
  3. ConfluentReaderSpec — the two unframed-refusal examples (`reader` /
     `recordReader`) become one over a shared predicate, keeping both payload
     shapes; the MOVED-field and PROMOTED-field examples become one over a
     shared `readFramed` rig.
  4. AvroWriteCorrectnessSpec — the three `.each.union[Branch]` examples
     (blocked framing, canonical framing, record face) were one fixture and one
     optic handed to three carriers: now one example, every assertion kept.
  5. AvroWriteCorrectnessSpec — the two `.each.fields` examples likewise
     differed only in how the same basket was framed.
  6. AvroJsonSpec — the enum/bytes/fixed RENDER example and the enum/bytes/fixed
     ROUND-TRIP example built two near-identical three-leaf schemas. One schema,
     one value, and the render half now asserts the WHOLE document
     (`avroToJson(record) === json`) instead of three field lookups — stronger
     than what it replaced.

MEASURED, not assumed. `project avroIntegration; stryker` re-run and diffed
against the pre-consolidation report by the full (file, line, column, mutator,
replacement) key: 431 Killed / 38 Survived / 17 NoCoverage / 6 CompileError /
122 Ignored BEFORE, and exactly the same after, with zero keys added, zero
removed, zero Killed -> anything. Mutation score unchanged at 88.68% total /
91.90% of covered. Ran twice during the phase (after folds 1-3 and after 4-6);
both diffs clean. Against this run's original baseline the whole branch is
+40 kills with zero regressions.

No assertion was weakened and no fixture dropped: every fold keeps every
`must`/`===` it merged, and three of them add one.

avro suite: 177 -> 159 `>>` registrations, 5,201 -> 5,122 test LOC. Against the
run's starting point (173 / 4,952) that is 14 FEWER examples for +40 mutant
kills; the +170 lines are the net of 242 lines of new high-leverage tests minus
79 removed here, i.e. 4.25 net lines per additional kill.

Gates on JDK 25: `avroIntegration/test` Passed 198, Failed 0, Errors 0;
`scalafixAll; scalafmtAll` then `scalafmtCheckAll; scalafixAll --check` clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
…Walk

#107 landed while this branch was open and rewrote the code the spec targets:
`sameFieldName` and the fuzzy per-field scan are gone, replaced by
`normalisedName` plus a weak-keyed `normalisedNameIndex` cache. The spec itself
needed no change — it was written against the public seam `AvroWalk.fieldNameAt`
precisely so it would survive an internals rewrite, and it agrees with the new
implementation cell for cell — but its `covers:` block named deleted symbols and
pre-#107 line numbers.

Re-aims it at the surviving constructs by name and drops the line numbers, and
states the division of labour with #107's own `NominalResolutionOracle`: that
one is the OLD ALGORITHM frozen verbatim (a refactor gate — "the new rung agrees
with the one it replaced"), this one is written from the DOCTRINE without
reading the implementation. Parity cannot detect a wrong doctrine, which is the
#104 failure mode; the doctrine oracle still passing against wholly rewritten
internals is the evidence that #107 preserved the contract and not just the code
path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
@kryptt
kryptt force-pushed the test/avro-mutant-kills branch from 844748d to cb4edec Compare September 18, 2026 16:29
@kryptt
kryptt merged commit 7ab894d into main Sep 18, 2026
20 of 23 checks passed
@kryptt
kryptt deleted the test/avro-mutant-kills branch September 18, 2026 20:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant