Skip to content

fix(avro): resolve field names by NAME when the codec proves it, and reject nested selectors (#95) - #98

Merged
kryptt merged 6 commits into
fix/fields-namedtuple-tuplen-96from
fix/nominal-field-resolution-hazard
Sep 18, 2026
Merged

kryptt merged 6 commits into
fix/fields-namedtuple-tuplen-96from
fix/nominal-field-resolution-hazard

Conversation

@kryptt

@kryptt kryptt commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #97 (fix/fields-namedtuple-tuplen-96). Merge that one first; this PR's base is
its branch, so the diff shown here is only my six commits.

Fixes the silent field mis-targeting surfaced while investigating #95, plus a second, independent
hazard with the same symptom found while reproducing it.

This does not close #95. That issue asks for an efficient whole-record construction primitive,
and nothing here addresses it. What it does close is the defect its reporter's codec shape exposed:
their schema carries computed/derived columns ("no equivalent of a computed/derived field mechanism"
in the issue body), which is exactly the shape that made every .field(_.x) past the divergence
read and write the wrong slot.

The hazard

.field(_.x) resolved x to the schema field at its DECLARATION INDEX and nothing else.
AvroWalk.fieldNameAt's whole body was fields.get(declIdx).name, guarded only against a
non-record parent (AvroWalk.scala:414) and an index past the end (AvroWalk.scala:422). The
field's name, its type and the record's arity were never consulted.

That is sound exactly while the codec's schema is positionally 1:1 with the case class — true by
construction for every kindlings-derived codec, and not true for a hand-written or
vulcan.Codec field list, which can add a computed column, drop one, or reorder:

// vulcan schema {a, computed, b, c}   vs   case class Three(a, b, c)
codecPrism[Three].field(_.b).getOption(bytes)      // Some("A0|B0")  <- read `computed`
codecPrism[Three].field(_.b).replace("B1")(bytes)  // rewrote `computed`

Valid Avro bytes, wrong content. Nothing catches it: get-put, put-get, put-put and modify-fusion
all hold on a mis-targeted optic (it is a perfectly lawful Optional onto the wrong field),
round-trips hold, and every fixture in AvroSpecFixtures is kindlings-derived, hence 1:1, hence
blind. One function, six call sites:

# Site file:line (before)
S1 AvroPrism.widenPath ← .field(_.x) AvroPrism.scala:249
S2 same ← selectDynamic AvroPrismMacro.scala:40
S3 AvroPrism.resolveFieldNames ← .fields(...) AvroPrism.scala:302
S4 AvroTraversal.widenSuffix ← .each.field(_.x) AvroTraversal.scala:158
S5 AvroTraversal.toFieldsTraversal ← .each.fields(...) AvroTraversal.scala:219
S6 every deeper hop after .union[B] / a nested .field AvroWalk.scala:351-380

The mechanism, and why all-or-nothing

One rung above position, firing only when the codec has proved the whole correspondence:

  1. Nominal, all-or-nothing. If EVERY case field maps to a DISTINCT schema field — exactly, or
    uniquely up to _/-/. and case — the map is total and injective, so it is the codec's own
    answer, not a guess. Partial or colliding coverage disqualifies the rung for every field.
  2. Positional (Field navigation must honor the schema's field names, not the Scala field names #35), unchanged. Where a name transform lands, because a transform REMOVES the
    literal Scala name by construction, so rung 1 cannot have fired.

The totality requirement is the load-bearing part, and it is measured, not argued. The per-field
form of rung 1 ("does THIS field's name appear?") INTRODUCES corruption: FpVisit(userId, user)
against legacy columns {uid, user_id} is correct by position today, and a per-field rung re-aims
userId at user_id — a currently-working call site silently moved to the wrong column. Requiring
totality makes the rung abstain there (user matches nothing) and position stays right.

An arity gate was rejected on the same kind of evidence: it refuses 13 of 28 legitimate call
sites — abbreviated wire names plus a trailing ingested_at, a dropped cachedHash, a trailing
checksum on an element codec, a digest in a nested codec — every one correct today because the
extra field is at the tail and shifts nothing. It buys 3 cells.

False-positive scorecard (ResolutionFalsePositiveSpec, 28 legitimate call sites)

Scored against SLOT TRUTH — which schema slot a write actually touched, never copy(...), which
cannot tell a wrong slot from a stale derived one.

CORRECT LOUD-REFUSAL SILENT-WRONG SILENT-MISS
before (published behaviour) 28 0 0 0
after (this PR) 27 1 0 0

The single moved cell is hatch-record-probe-absent — .fieldNamed on a name the reader schema
does not carry, None before and a construction-time refusal now. That is the deliberate change
below. The two TRIPWIRE: examples (zero SILENT-WRONG, zero SILENT-MISS) pass both ways and are
the whole safety argument for touching the resolver.

AvroFieldNamingSpec's other 8 examples — issue #35's own snake_case regression suite — pass
untouched.

.fieldNamed is checked at construction

The escape hatch every error message points at appended the literal with no schema lookup at
all
, although the prism already holds the schema. A typo — or the Scala field name passed where
the schema name was meant — was a runtime SILENT MISS: reads None, writes hand back the payload
unchanged and report success. Sending someone from a loud construction failure to a silent runtime
one is the wrong direction.

Two carve-outs: a MAP parent (.fieldNamed is also how a map KEY is addressed, keys are data,
and feature-detecting an absent key is legitimate — three cells pin it) and an unresolvable parent
path (the walk already reports that, and refusing would change the meaning of a prism deliberately
built against a drifted root schema). Record-level feature detection moves to
codec.schema.getField(name).

Two existing examples change: AvroFieldNamingSpec's "a bad explicit .fieldNamed misses (None),
it does not corrupt" pinned the defect and now asserts the refusal, and AvroBytesSpec keeps its
PathMissing coverage by storing the PathStep.Field("name") through the internal constructor that
example already uses two lines above.

The nested-selector hole (second, independent bug)

MacroSelectors.extractFieldName matched Lambda(_, Select(_, name)) with any receiver. A
nested path _.inner.y is Select(Select(Ident(_), "inner"), "y"), so it matched, yielded the bare
name "y", and every cursor macro resolved y on the parent:

// NOuter(inner: NInner(x, y), y)
codecPrism[NOuter].field(_.inner.y).getOption(bytes)  // Some("OUTER_Y")
codecPrism[NOuter].field(_.inner).field(_.y)          // Some("INNER_Y")

Silent corruption on a perfectly 1:1, kindlings-derived codec, with no schema divergence involved
at all
— and the macros' own "nested paths are not yet supported inside a single call; chain them"
abort was unreachable for exactly the shape it was written for. 100% decidable at compile time.
LensMacro never had this (it uses the strict extractSingleFieldName plus a knownFields.contains
check); the cursor macros simply never got the same treatment.

Both halves land: the extractor now requires the Select receiver to be the lambda parameter
(keeping the wrapper tolerance that is the only reason it exists beside the strict form), and
requireCaseField[A] aborts when a single-hop selector names a non-case-field — those used to pass
through as a literal field name and miss at runtime. Five call sites: AvroPrismMacro.fieldImpl /
fieldTraversalImpl, JsonPrismMacro.fieldImpl / fieldTraversalImpl (eo-circe),
JsoniterPrismMacro.fieldName (eo-jsoniter, shared by prism and traversal). grep finds zero
nested-selector call sites in the repo; the circe, jsoniter, avro and generics suites all pass.

Documented residual limitations (ResolutionResidualSpec, pinned as executable examples)

The memo this PR implements asserted these by inspecting the rule. They are now run, and every
one behaves as predicted:

Cause 1 — no name signal (names unrecoverable AND the list permuted, so the rung abstains and
position decides, wrongly). Reachable only by a differential probe of the codec or an explicit
declaration.

  • 5a {beta_name, alpha_name} vs {alpha, beta} — reads beta for alpha
  • b2 {seq_no, event_ts} vs {occurredAt, seqNo} — reads the sequence number; a
    timestamp-millis logical type annotates the same physical long, so it cannot disambiguate
  • b6 compensating arity (one case field dropped, one computed column added): counts equal, position
    wrong from the insertion point on
  • at3 userId_ and user_id both normalise to userid — ambiguity is no signal, so position
    lands on the stale duplicate

Cause 2 — a MISLEADING name signal (a column bears a case field's name but holds a different
value; the map is total and injective, so it is trusted).

  • at1 {digest, id, raw_ident, payload} — id is a derived public identifier
  • at2 {USER_ID, user_ident, balance} — USER_ID is a stale legacy column

Both cause-2 cells were wrong on 0.15.1 too, with one real cost: the corrupt value is now more
plausible (pub-REAL-ID rather than a digest length). at1's payload is fixed by the change
— it read id before. .fieldNamed reaches all six, and that is asserted too.

Accepted behaviour change: a hand-written codec that permutes the Scala names (writes case
field a into a schema field literally named b, and vice versa) resolved correctly by position and
now resolves by name, i.e. wrongly. No name transform can produce that shape — a transform is a
function of the name alone — but a hand-written field list can.

Cost

Zero per operation, verified from bytecode rather than asserted. javap -p -c over the compiled
avro classes puts every reference to resolveFieldName / fieldNameAt / requireField* in
AvroPrism's widenPath / widenPathNamed / toFieldsPrism / resolveFieldNames and
AvroTraversal's widenSuffix / widenSuffixNamed / toFieldsTraversal. AvroFocus,
AvroFocus$Leaf, AvroFocus$Fields, AvroRecordPrism, AvroRecordTraversal and
AvroBinaryCursor contain none.

Construction time is not free, and the number is reported rather than waved at (best of 5 ×
200 000 builds, same box, same session):

base this PR
identity names .field(_.name) 84.9 98.6 ns/build
snake_case .field(_.clickId) 85.1 264.0 ns/build
nested .field(_.meta).field(_.performanceSourceId) 116.9 430.3 ns/build

The gap is inherent to totality: the rung resolves every case field before trusting any one of
them, so a snake_case parent runs a normalised name compare per (case field × schema field) pair
where position ran none. A first cut using Set[Int] for injectivity cost 316 / 462; the shipped
version scans the filled prefix of the result array instead and adds an early "more case fields than
schema fields ⇒ no total map" precondition. Paid once per drilled prism, strictly off the read/write
path.

Back-compat

Every touched signature is private[avro], but transparent inline bakes the accessor into
caller bytecode (inline$widenPath$i1 gains a List[String] parameter), so downstream must
recompile, not re-jar. MiMa is off build-wide (tlMimaPreviousVersions := Set.empty), so this is
a release note, not a build gate; mimaReportBinaryIssues is green.

Gates

sbt "scalafixAll; scalafmtAll", then — as the last step, after every source edit —
scalafmtCheckAll + scalafmtSbtCheck + benchmarks/scalafmtCheck + scalafixAll --check: all
[success]. avroIntegration/test 212/212, root test all green (409 in tests, 212 avro, 92
core, 56 jsoniter, 42 circe, …), mimaReportBinaryIssues clean, docs/mdoc + docs/laikaSite
build.

Follow-ups NOT in this PR

  • TargetingLaws / TargetingTests in cats-eo-laws — an optic vs. an extrinsic reference
    accessor. The only thing that can catch a resolution mechanism's own mistake, and the hazard
    class is carrier-wide (circe/jsoniter resolve by the literal Scala name, zio.schema's
    EoAccessorBuilder and kyo's Record.lens[F]("name") are the same shape). The fixtures land here;
    the reusable ruleset is next.
  • A differential codec probe — encode two sentinels through codec.encode, diff the slots,
    observe which schema field a case field actually lands in. The only mechanism that reaches residual
    cause 1, needs no declaration and no name heuristics, and would turn a wrong hand-written field
    mapping from believed into refused. Must degrade to the ladder rather than refuse.
  • AvroCodec.parseInputUnsafe's fabricated empty record (AvroCodec.scala:329-339) — a parse
    failure yields new GenericData.Record(schema), so modifyUnsafe on 4 junk bytes returns
    {"items": null}: a valid payload with all data erased. Adjacent silent-corruption hole, unrelated
    to resolution, cheap to fix.
  • AvroVulcan.codec's eager val schema; .at(i) has no ARRAY check (silent Miss on the whole byte
    face); AvroVulcan.codecMapped as a user-side escape for residual cause 1.

🤖 Generated with Claude Code

https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

kryptt and others added 6 commits September 18, 2026 10:09
…s it (#95)

`.field(_.x)` resolved `x` to the schema field at its DECLARATION INDEX and
nothing else: `AvroWalk.fieldNameAt`'s whole body was `fields.get(declIdx).name`,
guarded only against a non-record parent and an index past the end. The field's
NAME, its TYPE and the record's ARITY were never consulted.

That is sound exactly while the codec's schema is positionally 1:1 with the case
class — true by construction for every kindlings-derived codec, and NOT true for
a hand-written or `vulcan.Codec` field list, which can add a computed column,
drop one, or reorder. On those the optic silently targets the WRONG SLOT:

    schema {a, computed, b, c}  vs  case class Three(a, b, c)
    codecPrism[Three].field(_.b).getOption(bytes)     // Some("A0|B0")  <- `computed`
    codecPrism[Three].field(_.b).replace("B1")(bytes) // rewrites `computed`

and produces valid Avro bytes with wrong content. Nothing catches it: get-put,
put-get, put-put and modify-fusion all HOLD on a mis-targeted optic (it is a
perfectly lawful Optional onto the wrong field), round-trips hold, and every
fixture in `AvroSpecFixtures` is kindlings-derived, hence 1:1, hence blind.
One function, six call sites: `.field`, `selectDynamic`, `.fields`,
`.each.field`, `.each.fields`, and every deeper hop after a `.union[B]`.

The fix adds one rung ABOVE position, and only fires when the codec has proved
the whole correspondence:

  1. NOMINAL, all-or-nothing. If EVERY case field maps to a DISTINCT schema
     field — exactly, or uniquely up to `_`/`-`/`.` and case — the map is total
     and injective, so it is the codec's own answer, not a guess.
  2. POSITIONAL (issue #35), unchanged. This is where a name transform lands,
     because a transform REMOVES the literal Scala name by construction.

All-or-nothing is the load-bearing part, and it is measured, not assumed. The
per-field form of rung 1 ("does THIS field's name appear?") INTRODUCES
corruption: `FpVisit(userId, user)` against legacy columns `{uid, user_id}` is
correct by position today, and a per-field rung re-aims `userId` at `user_id`.
Requiring totality makes the rung abstain there — `user` matches nothing — and
position stays right. An arity gate was rejected on the same evidence: it
refuses 13 of 28 legitimate call sites (a trailing `ingested_at`, a dropped
`cachedHash`, a trailing checksum) to buy 3 cells.

Per-operation cost is zero: resolution runs once, at prism construction, off
the cached schema. `to`, `from`, `scan`, `spliceAff`, `spliceFoci` and both
erased bridges contain no reference to it on either carrier.

Known residual, pinned as specs rather than prose: a codec that both renames
beyond recognition AND reorders (equal arity, no name hit) still resolves by
position and is still wrong. Known behaviour change: a codec that PERMUTES the
Scala names was right by position and is now wrong by name — no transform can
produce that shape, only a hand-written field list.

Back-compat: every touched signature is `private[avro]`, but `transparent
inline` bakes the accessor into CALLER bytecode, so downstream must recompile,
not re-jar.

Tests:
- `DivergentCodecs` (new) — hand-written vulcan codecs whose schema is not
  positionally 1:1. The standing gap: every existing fixture is derived.
- `AvroNominalResolutionSpec` (new) — the 9 repro cells (computed field on the
  byte and record faces, the leaf past the divergence, `.fields` grouped
  read+write, `.each.field`, a nested record's own divergence, a reversed field
  list, snake_case + a computed field) plus the two false-positive controls.
  Every repro cell failed before this change; both controls passed.
- `ResolutionFalsePositiveSpec` (new) — 28 legitimate, currently-working call
  sites scored against SLOT TRUTH. 28/28 CORRECT before and after: that
  scorecard not moving is the entire safety argument for the mechanism.
- `AvroWalkSpec` — three direct `resolveFieldName` / `fieldNameAt` calls take
  the new case-name list (mechanical, `Nil`).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
…misses

The scaladoc that shipped with issue #35 was wrong in three ways, and each one
is a reason the hazard in #95 went unnoticed for a release:

- It warned only about schema field ORDER divergence. Order is not the shape
  that bit: a COMPUTED/derived schema column with no case-class parameter (or a
  dropped one) shifts every later slot on a codec whose field order is perfectly
  sensible.
- It never said the rule governs `.fields`, `selectDynamic` or `.each.field`.
  They share one resolver, and a grouped `.fields` write on a divergent parent
  damaged two slots and silently no-op'd a third.
- It advertised kebab-case as a supported transform. Kebab cannot produce a
  legal Avro schema at all — `SchemaParseException: Illegal character in:
  click-id` — so the claim was never true.

Rewritten to state the actual ladder (name-when-total, else position), the
precondition the positional rung needs, the three shapes that stay wrong after
it, and the one accepted behaviour change. Same pass over the docs site's
"Field navigation is by SCHEMA name" section and `AvroWalk`'s internal banner,
plus a CHANGELOG entry carrying the recompile-not-rejar note.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
Three codec shapes are still resolved to the wrong schema field after the
nominal rung, and the decision memo asserted them by INSPECTING the rule rather
than running it. This runs them. All eight cells behave exactly as predicted,
so the ledger below is measured, not argued:

  cause 1, no name signal (the names are unrecoverable AND the list is permuted,
  so the rung abstains and position decides, wrongly)
    5a  {beta_name, alpha_name} vs {alpha, beta}    -> reads beta for alpha
    b2  {seq_no, event_ts} vs {occurredAt, seqNo}   -> reads the sequence number
        (a timestamp-millis logical type annotates the same physical long, so
        it cannot disambiguate either)
    b6  compensating arity: one case field dropped, one computed column added,
        counts equal, position wrong from the insertion point on
    at3 `userId_` and `user_id` both normalise to `userid` — ambiguity is no
        signal, so the rung disqualifies itself and position lands on the stale
        duplicate

  cause 2, a MISLEADING name signal (a column BEARS a case field's name but
  HOLDS a different value; the map is total and injective, so it is trusted)
    at1 {digest, id, raw_ident, payload} — `id` is a derived public identifier
    at2 {USER_ID, user_ident, balance} — `USER_ID` is a stale legacy column

Both cause-2 cells were wrong on 0.15.1 too (they read `digest` / position 0),
with one real cost: the corrupt value is now more PLAUSIBLE, `pub-REAL-ID`
rather than a digest length. Note at1's `payload` is FIXED by the change — it
read `id` before.

Also pins the accepted regression (a PERMUTING rename was right by position and
is now wrong by name) and asserts that `.fieldNamed` reaches every residual,
which is what the error messages and the docs both point at.

Reaching cause 1 needs a differential probe of the codec itself (encode
sentinels, observe which slot they land in) or an explicit declaration; no
amount of name matching helps. Filed as a follow-up, not attempted here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
`.fieldNamed("schema_name")` appended the literal to the path with NO schema
lookup at all — `widenPathNamed` was one `widenPathStep(PathStep.Field(name))` —
although the prism already holds the schema. A typo, or the Scala field name
passed where the schema name was meant, therefore produced a runtime SILENT
MISS: reads return `None`, writes hand back the payload unchanged and report
success.

That is the same failure class the hatch exists to avoid, and it is the failure
class every error message in this resolver points AT: "navigate by explicit
schema name with .fieldNamed(...)". Sending someone from a loud construction
failure to a silent runtime one is the wrong direction.

Two deliberate carve-outs:
- a MAP parent. `.fieldNamed` is also how a map KEY is addressed, keys are data
  rather than schema fields, and feature-detecting an absent key is legitimate
  (`hatch-map-key-present` / `-absent` / `-write` in the false-positive spec
  pin all three).
- an unresolvable parent path. The walk already reports that at runtime, and
  refusing here would change the meaning of a prism deliberately built against
  a drifted root schema.

Record-level feature detection ("does this schema carry X?") moves to the
schema, where it belongs: `codec.schema.getField(name)`.

Two existing examples change, both deliberately:
- `AvroFieldNamingSpec`'s "a bad explicit .fieldNamed misses (None), it does
  not corrupt" PINNED the defect. Rewritten to assert the refusal.
- `AvroBytesSpec` used `.fieldNamed("name")` on a deliberately narrowed root
  schema to reach the walker's `PathMissing` arm. The coverage is kept by
  storing the `PathStep.Field("name")` through the internal constructor that
  example already uses two lines above.

Also amends the normative laws bullet, which promised "structurally drifted
payloads Miss silently" without distinguishing PAYLOAD drift (still a runtime
miss, still undetectable) from a name absent from the READER schema (now a
construction-time refusal on both `.field` and `.fieldNamed`).

The false-positive scorecard moves by exactly one cell of 28 —
`hatch-record-probe-absent`, None -> refusal — and by no other.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
#95)

`MacroSelectors.extractFieldName` — the selector parser shared by every cursor
macro — matched `Lambda(_, Select(_, name))` with ANY receiver. A nested path
`_.inner.y` is `Select(Select(Ident(_), "inner"), "y")`, so it matched, yielded
the bare name `"y"`, and the macro resolved `y` on the PARENT record.

Where the parent carries a field of that name — and a record holding a nested
record often does — that is not a miss. It is a well-typed, perfectly lawful
optic aimed at the wrong field:

    NOuter(inner: NInner(x, y), y)
    codecPrism[NOuter].field(_.inner.y).getOption(bytes)   // Some("OUTER_Y")
    codecPrism[NOuter].field(_.inner).field(_.y)           // Some("INNER_Y")

Silent corruption on a perfectly 1:1, derived codec, with no schema divergence
involved at all — a second, independent hazard from the resolution one, and
100% decidable at compile time. The macros' own "nested paths are not yet
supported inside a single call; chain them" abort was UNREACHABLE for exactly
the shape it was written for.

`LensMacro` never had this: it uses the strict `extractSingleFieldName` (which
requires the `Select` receiver to be the lambda parameter) plus a
`knownFields.contains` check. The cursor macros simply never got the same
treatment. Both halves land here:

- `extractFieldName` now unwraps `Inlined` / `Typed` around the lambda, around
  its `Select` body AND around the `Select`'s RECEIVER, then requires that
  receiver to be an `Ident`. That keeps the wrapper tolerance the cursor macros
  need (which is the only reason this function exists beside the strict one)
  and drops the receiver looseness, which was never intentional.
- `requireCaseField[A]` aborts when a single-hop selector names something that
  is not a case field of the parent. Those used to be passed through as a
  literal field name and miss at runtime; the declaration index comes back `-1`
  for them, which is also the legitimate "NamedTuple parent" signal, so the two
  cannot be told apart downstream. They are told apart here. Skipped when `A`
  has no case fields at all, which is the shape the literal fallback exists for.

Five call sites, all covered: `AvroPrismMacro.fieldImpl` / `fieldTraversalImpl`,
`JsonPrismMacro.fieldImpl` / `fieldTraversalImpl` (eo-circe), and
`JsoniterPrismMacro.fieldName` (eo-jsoniter, shared by prism and traversal).
The two traversal messages also gained the "chain them" hint their prism twins
already carried.

`grep` finds zero nested-selector call sites in the repo, and the circe,
jsoniter, avro and generics suites all pass untouched — this closes a hole, it
does not move any working code.

Tests: `NestedSelectorMacroErrorSpec` in avro, circe and jsoniter — the nested
prism selector, the nested traversal selector, the not-a-case-field selector,
and a runtime row asserting the CHAINED form reaches the inner field. Every
compile-error row produced NO error at all before this change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
Resolution is construction-time only, but it is not free, and the first cut
paid more than it needed to. Measured with the prototype's own harness (best of
5 x 200_000 builds, same box, same session, `ConstructionCostProbe`):

                                     base    Set[Int]   this
  identity  .field(_.name)          84.9     115.2      98.6  ns/build
  snake     .field(_.clickId)       85.1     316.0     264.0  ns/build
  nested    .field(_.meta).field(…) 116.9    462.2     430.3  ns/build

Two changes, both free:

- `Set[Int]` -> a linear scan over the filled prefix of the result array. The
  set was allocating a boxed Integer and a new set node per case field, on a
  list that is a handful of entries; the scan is O(n^2) on an n that is the
  case-class arity.
- An early precondition: a total, injective map from case fields into schema
  fields cannot exist when the case class has MORE fields than the record, so
  that shape skips the scan entirely. Same answer, no work.

The remaining gap is inherent to the rule: the total-nominal rung resolves
EVERY case field before it trusts any one of them, so a snake_case parent runs
a normalised name compare per (case field x schema field) pair where position
ran none. That is the price of not re-aiming a working call site at a lucky
single match, it is paid once per drilled prism at construction, and it stays
strictly off the read/write path.

Zero per-operation cost re-verified from bytecode rather than asserted:
`javap -p -c` over the compiled avro classes puts every reference to
`resolveFieldName` / `fieldNameAt` / `requireField*` in `AvroPrism`'s
`widenPath` / `widenPathNamed` / `toFieldsPrism` / `resolveFieldNames` and
`AvroTraversal`'s `widenSuffix` / `widenSuffixNamed` / `toFieldsTraversal`.
`AvroFocus`, `AvroFocus$Leaf`, `AvroFocus$Fields`, `AvroRecordPrism`,
`AvroRecordTraversal` and `AvroBinaryCursor` contain none.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
@kryptt

kryptt commented Sep 18, 2026

Copy link
Copy Markdown
Contributor Author

⚠️ These six commits never reached main. This PR was merged into fix/fields-namedtuple-tuplen-96 28 seconds after that branch had itself been squash-merged to main as #97 (8d0cb29), so the merge landed on a ref that no longer had a path to main — confirmed three ways: git merge-base --is-ancestor 7fc97fe origin/main is false, caseNames (the nominal rung's own identifier) has grep -c = 0 in main:avro/.../AvroWalk.scala, and gh api .../compare/main...7fc97fe reports status: diverged.

GitHub refuses to retarget a merged PR ("Cannot change the base branch of a closed pull request"), so the content has been recovered onto a fresh branch off main with these six commits cherry-picked unchanged — the resulting diff is byte-identical to this one (2183 lines, verified with diff -q).

Review continues in #105 (#105); this PR stays closed as a record.

kryptt added a commit that referenced this pull request Sep 18, 2026
…field

The multiplier issue #103 set out to remove survived one level down.
`totalNominalIndex` asked `nominalIndex` once per case field, and every
`nominalIndex` that missed the exact-name hash called `normalisedNameIndex`,
which probes the cache with a freshly allocated `SchemaQuery` key. On the
transformed path — where EVERY exact lookup misses by construction, which is
the whole population the nominal rung exists for — that is `arity` ConcurrentHashMap
probes and `arity` throwaway keys per resolution, all of them answering with the
same map. At arity 66: 66 probes, 66 keys, one map.

The index is now threaded through the `@tailrec` loop as a recursion parameter:
`null` until some case field needs it, the resolved map afterwards. Threaded and
not hoisted above the loop, because hoisting would make the identity-named codec
build or fetch an index it never reads — that path must stay untouched, and it is:
the exact-name hash rung is still FIRST, still `record.getField(scalaName)`, still
allocation-free.

`nominalIndex` disappears as a separate method (it was the per-field wrapper), its
two rungs inlined into the loop where the index can be carried across iterations.
Verdicts are unchanged by construction — same rungs, same order, same `Ambiguous`
demotion — and `NominalResolutionParitySpec` re-confirms it: 151,515 synthetic
(schema, case-name list, declIdx) cells plus every real #98 fixture, zero
mismatches against the frozen pre-index oracle.

Measured with `com.sun.management.ThreadMXBean.getThreadAllocatedBytes`, warm,
bytes per resolution (the harness the PR's disclosed figures came from):

  shape        n=2    n=12    n=33    n=66     per case field
  identity      24      64     152     280     4 B  (the `out` array, unchanged)
  snake before 280   1,600   4,376   8,728     132 B
  snake after  256   1,336   3,608   7,168     108 B

The saving is exactly `(arity - 1) x 24 B` — one `SchemaQuery` per case field
beyond the first — confirming the key was NOT being scalar-replaced by escape
analysis. Net: -17.9% on the transformed path at n=66. What remains is the
normalised `String` itself, ~104 B per case field, which is the disclosed trade
and needs the cursor-hash probe to close. The identity-named path is
byte-identical before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
kryptt added a commit that referenced this pull request Sep 18, 2026
…dicts

RED for issue #103. Nothing here changes behaviour; it makes both halves of
the problem executable before the fix touches anything.

`NominalResolutionOracle` is a VERBATIM copy of PR #98's nominal rung. It is
the differential oracle: a cost rewrite of a resolution rule is only safe if
it is verdict-identical, and "the examples still pass" cannot show that — an
example pins one cell, while the rung's semantics are a function over
(schema, case-field list, declaration index).

`NominalResolutionParitySpec` diffs the live rung against that oracle over
259 synthetic schemas x 585 case-name lists (151,515 pairs, every declIdx in
[-1..arity]) built from deliberately confusable names — exact, case and
separator collisions — plus every codec behind the three #98 suites. It also
names each semantic separately, so a failure says WHICH rule broke:
exact-beats-normalised, two-normalised-hits-is-none, the `arity >
fields.size` short-circuit, whole-list injectivity, the empty-caseNames
(NamedTuple) abstention, and the out-of-range declIdx abstention.

`NominalResolutionCostSpec` is the failing half. It asserts nothing in
nanoseconds — this box is noisy and only within-run ratios are load-bearing
— but on the SHAPE: double the field count and a quadratic rung costs ~4x, a
linear one ~2x. Against the current implementation:

  snake_cased n=66  warm  472,474 ns per resolution
  2n/n = 3.97 (quadratic)      4n/n = 15.44 (quadratic)
  n=66 snake / identity = 242x

`totalNominalIndex` widens from `private` to `private[avro]` so the parity
spec can read its verdicts directly rather than inferring them from a
resolved field name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
kryptt added a commit that referenced this pull request Sep 18, 2026
…field

The multiplier issue #103 set out to remove survived one level down.
`totalNominalIndex` asked `nominalIndex` once per case field, and every
`nominalIndex` that missed the exact-name hash called `normalisedNameIndex`,
which probes the cache with a freshly allocated `SchemaQuery` key. On the
transformed path — where EVERY exact lookup misses by construction, which is
the whole population the nominal rung exists for — that is `arity` ConcurrentHashMap
probes and `arity` throwaway keys per resolution, all of them answering with the
same map. At arity 66: 66 probes, 66 keys, one map.

The index is now threaded through the `@tailrec` loop as a recursion parameter:
`null` until some case field needs it, the resolved map afterwards. Threaded and
not hoisted above the loop, because hoisting would make the identity-named codec
build or fetch an index it never reads — that path must stay untouched, and it is:
the exact-name hash rung is still FIRST, still `record.getField(scalaName)`, still
allocation-free.

`nominalIndex` disappears as a separate method (it was the per-field wrapper), its
two rungs inlined into the loop where the index can be carried across iterations.
Verdicts are unchanged by construction — same rungs, same order, same `Ambiguous`
demotion — and `NominalResolutionParitySpec` re-confirms it: 151,515 synthetic
(schema, case-name list, declIdx) cells plus every real #98 fixture, zero
mismatches against the frozen pre-index oracle.

Measured with `com.sun.management.ThreadMXBean.getThreadAllocatedBytes`, warm,
bytes per resolution (the harness the PR's disclosed figures came from):

  shape        n=2    n=12    n=33    n=66     per case field
  identity      24      64     152     280     4 B  (the `out` array, unchanged)
  snake before 280   1,600   4,376   8,728     132 B
  snake after  256   1,336   3,608   7,168     108 B

The saving is exactly `(arity - 1) x 24 B` — one `SchemaQuery` per case field
beyond the first — confirming the key was NOT being scalar-replaced by escape
analysis. Net: -17.9% on the transformed path at n=66. What remains is the
normalised `String` itself, ~104 B per case field, which is the disclosed trade
and needs the cursor-hash probe to close. The identity-named path is
byte-identical before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
kryptt added a commit that referenced this pull request Sep 18, 2026
…68 lines

PR #98 added an all-or-nothing nominal rung to `AvroWalk.fieldNameAt`: resolve
`.field(_.x)` by NAME when every case field maps to a distinct schema field,
otherwise abstain to position. Four clauses; only TOTALITY was pinned.
Injectivity, exact-beats-normalised, and ambiguity-implies-abstain had no
fixture that could tell a correct implementation from a broken one — the sole
ambiguity fixture (`ResolutionResidualSpec` "at3") passes for the wrong reason,
because an unrelated unmapped field forces totality-driven abstention first and
masks the ambiguity branch entirely.

That is not cosmetic. Issue #104 recorded an experimental alias rung silently
breaking a case the rung gets right while all 212 avro tests stayed green: a
MIS-TARGETED optic satisfies get-put, put-get, put-put and modify fusion
perfectly, so no optic law family can see this.

`AvroNominalDoctrineSpec` states the doctrine as a declarative oracle and checks
`fieldNameAt` against it EXHAUSTIVELY over ~66k cells: every injective field
list of length 1..3 from a 6-name pool (each name both a legal Avro field name
and a legal Scala identifier), crossed with every case-name list of length 0..3
WITH repetition, crossed with declIdx from -1 to arity+1. The oracle is written
from the doctrine, not copied from the implementation — parity against a frozen
copy of the code cannot detect a wrong doctrine, which is precisely the #104
failure mode.

Enumeration, not `Gen`: the discriminating cells are threshold cells a default
generator would essentially never produce — a duplicate resolution landing on
slot 0, a fuzzy ambiguity whose FIRST hit is at index 0 versus at index >= 1,
and `declIdx` exactly equal to the schema's field count.

Verified by applying each mutant to `AvroWalk.scala` and confirming the property
goes red: 291 (exact fast path), 301 / 304 (the injectivity scan), 169 / 171 /
173 (ambiguity => abstain), 213 (`sameFieldName`'s char inequality), 310
(`declIdx >= fields.size` at the exact boundary), and 242 / 244 / 245 / 249 (the
four disjuncts of the all-or-nothing precondition).

Deletes the subsumed `"fieldNameAt: non-record parent and out-of-range declIdx
both throw loudly"` example: it probed `declIdx = 99` against a 2-field schema,
a cell where `>=` and a `>` mutant agree, so it pinned neither boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
kryptt added a commit that referenced this pull request Sep 18, 2026
…ld (#103) (#107)

* test(avro): pin the nominal rung's quadratic cost, and freeze its verdicts

RED for issue #103. Nothing here changes behaviour; it makes both halves of
the problem executable before the fix touches anything.

`NominalResolutionOracle` is a VERBATIM copy of PR #98's nominal rung. It is
the differential oracle: a cost rewrite of a resolution rule is only safe if
it is verdict-identical, and "the examples still pass" cannot show that — an
example pins one cell, while the rung's semantics are a function over
(schema, case-field list, declaration index).

`NominalResolutionParitySpec` diffs the live rung against that oracle over
259 synthetic schemas x 585 case-name lists (151,515 pairs, every declIdx in
[-1..arity]) built from deliberately confusable names — exact, case and
separator collisions — plus every codec behind the three #98 suites. It also
names each semantic separately, so a failure says WHICH rule broke:
exact-beats-normalised, two-normalised-hits-is-none, the `arity >
fields.size` short-circuit, whole-list injectivity, the empty-caseNames
(NamedTuple) abstention, and the out-of-range declIdx abstention.

`NominalResolutionCostSpec` is the failing half. It asserts nothing in
nanoseconds — this box is noisy and only within-run ratios are load-bearing
— but on the SHAPE: double the field count and a quadratic rung costs ~4x, a
linear one ~2x. Against the current implementation:

  snake_cased n=66  warm  472,474 ns per resolution
  2n/n = 3.97 (quadratic)      4n/n = 15.44 (quadratic)
  n=66 snake / identity = 242x

`totalNominalIndex` widens from `private` to `private[avro]` so the parity
spec can read its verdicts directly rather than inferring them from a
resolved field name.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* perf(avro): one cached name index per schema, not a scan per case field (#103)

GREEN for issue #103. The nominal rung's doctrine is unchanged — total and
injective or abstain entirely, exact before normalised, two normalised hits
is no hit. Only its cost moves.

`nominalIndex` tried Avro's exact-name hash and, on a miss, linear-scanned
every schema field calling `sameFieldName`, deliberately without an early
exit: it has to see a SECOND match before it can call a name ambiguous. Run
once per case field, that is `arity x fields.size x nameLength` per drilled
hop — and it is charged to exactly the codecs the rung exists for, because a
name transform (`withSnakeCaseFieldNames`, a custom `transformFieldNames`, a
`vulcan.Codec` rename map) GUARANTEES the exact-name miss. An identity-named
codec answers from the hash and never scanned, so the population the rung was
written for was the only one paying.

Now: one O(n) pass over the record's fields builds a normalised-name ->
position index, and each case field is one lookup. Collision detection is
preserved and merely MOVED — a key already present while BUILDING is the same
signal the scan produced on its second hit, seen once instead of per query.
Normalisation is per-CHARACTER (`Character.toLowerCase`), not
`String.toLowerCase`, which is locale-sensitive and can change a string's
length; that keeps the key exactly as discriminating as the cursor walk it
replaces.

The exact-name path is untouched and still allocation-free, so the
identity-named codec neither builds nor consults an index.

Cache: a `ConcurrentHashMap` side table keyed on the schema's REFERENCE
identity behind WEAK keys.
  - identity, not `Schema` itself: `Schema.equals` is deep-structural and
    `Schema` is mutable (`addProp` resets its cached hash), so a structural
    key would pay a deep compare per lookup and could lose an entry under
    annotation;
  - not Avro's own per-schema props, the obvious slot: a prop is part of
    `equals`, of `toString` and of the JSON and canonical forms, so a cache
    parked there would change what the schema SERIALISES AS;
  - weak, not bounded: schemas are usually module-level and immortal, where
    both are equivalent, but a registry client resolving writer schemas per
    message would leak under a strong map and thrash under a bound. The value
    holds no reference back to the schema, so the entry is genuinely
    collectable; `NominalNameIndexCacheSpec` proves it by dropping a schema
    and watching it be collected.

Measured, ns per resolution, MIN of 25 trials with both implementations timed
in the same JVM run (the frozen pre-index rung is kept in the test tree as
`NominalResolutionOracle`); schema construction is outside the clock:

  identity     n= 2   warm       34 ->       36     cold      108 ->      107
  identity     n=12   warm      194 ->      196     cold      289 ->      256
  identity     n=66   warm    2,128 ->    2,122     cold    5,902 ->    4,067
  snake_cased  n= 2   warm      490 ->      224     cold      605 ->      757
  snake_cased  n=12   warm   17,271 ->    1,361     cold   18,367 ->    3,963
  snake_cased  n=66   warm  496,939 ->    8,745     cold  500,700 ->   20,466

  2n/n  3.97 -> 2.19        4n/n  15.44 -> 5.12
  n=66 snake / identity  242x -> 3.9x

warm = repeated resolution against one schema, which is what a `def`-shaped
or parameterised-`given` optic pays per operation. cold = first resolution
against each of 200 never-before-seen schemas, index build included.

The cold column REGRESSES on small records, and says so: building the index
costs more than the scan it replaces when there is almost nothing to scan.

  cold, snake_cased, med of 25 trials (two runs)
  n=1   147 ->  434 / 148 ->  395    +287 / +247 ns   2.9x  SLOWER
  n=2   719 ->  831 / 723 ->  765    +113 /  +41 ns   1.2x  SLOWER
  n=3  1339 -> 1126 / 1320 -> 1169    -213 / -151 ns  0.85x
  n=4  2296 -> 1538 / 2294 -> 1518    -758 / -776 ns  0.67x

Crossover at n=3; the penalty is a few hundred nanoseconds, once per schema,
amortised over every later hop into it.

And it trades bytes for time on the transformed path, which the scan did not:
`normalisedName` materialises a String where `sameFieldName` walked two
cursors, so a miss now allocates ~128 B per case field (deterministic, via
`getThreadAllocatedBytes`):

  identity     n= 2     24 ->     24 B      n=12    64 ->    64 B
  identity     n=66    280 ->    280 B
  snake_cased  n= 2     24 ->    280 B      n=12    64 -> 1,600 B
  snake_cased  n=66    280 ->  8,728 B

Per OPERATION, not per construction, for a `def`-shaped optic. ~490 us of CPU
for ~8 KB of nursery at n=66 is the right way round, but "B/op gates, ns/op
advises" is this repo's doctrine, so it is on the record rather than implied.
Hoisting the index lookup out of the per-case-field loop does not close it:
the normalised String is ~104 B of the 128 and would survive untouched, and
that hoist was built and measured at byte-for-byte unchanged. Closing it
means probing the index by a cursor-computed hash of the normalised name and
confirming the candidate with the allocation-free pairwise comparison — a
larger change than #103 asked for.

The 151,515-cell verdict diff against the frozen implementation is EMPTY.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* docs(avro): the residual is unreachable from the schema, only from a probe

Three changes, all documentation.

The nominal rung's known-residual note now says what CANNOT close the
residual, not only that it is open. A 31-agent investigation settled this
today and the idea keeps being re-proposed: neither the record's schema NAME
nor its parsing FINGERPRINT can detect the residual pair (two renamed,
swapped columns of the same physical type within one record), because both
candidate readings carry the same fullname and the same fingerprint by
construction, and a schema-derived value only ever sees ONE of the two
operands a comparison would need. The only mechanism that reaches the class
is the differential codec probe — issue #100.

The pre-index's two COSTS are now on the record, in the CHANGELOG and beside
the code that pays them, because the entry as written read as if nothing
regressed:

  - `normalisedName` materialises a String, so a resolution that misses the
    exact-name hash allocates ~128 B per case field where the pairwise cursor
    walk it replaced allocated nothing: 24 -> 280 B at n=2, 280 -> 8,728 B at
    n=66 per resolution (deterministic, via getThreadAllocatedBytes). The
    identity-named path is byte-identical, as claimed. This is per OPERATION
    for a `def`-shaped optic, so it is nursery traffic on the hot path, not
    construction-time only — and "B/op gates, ns/op advises" is the repo's
    own doctrine, so a future gate will see it.
  - Building the index costs more than the scan it replaces on a one- or
    two-field record, so the FIRST resolution against one is slower. The
    crossover is at n=3.

Both are the right trade — ~490 us of CPU per resolution traded for ~8 KB of
nursery at n=66, and a few hundred nanoseconds once per small schema — but
they belong in the record rather than in a test comment.

Plus the CHANGELOG entry for #103.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* perf(avro): resolve the name index once per resolution, not per case field

The multiplier issue #103 set out to remove survived one level down.
`totalNominalIndex` asked `nominalIndex` once per case field, and every
`nominalIndex` that missed the exact-name hash called `normalisedNameIndex`,
which probes the cache with a freshly allocated `SchemaQuery` key. On the
transformed path — where EVERY exact lookup misses by construction, which is
the whole population the nominal rung exists for — that is `arity` ConcurrentHashMap
probes and `arity` throwaway keys per resolution, all of them answering with the
same map. At arity 66: 66 probes, 66 keys, one map.

The index is now threaded through the `@tailrec` loop as a recursion parameter:
`null` until some case field needs it, the resolved map afterwards. Threaded and
not hoisted above the loop, because hoisting would make the identity-named codec
build or fetch an index it never reads — that path must stay untouched, and it is:
the exact-name hash rung is still FIRST, still `record.getField(scalaName)`, still
allocation-free.

`nominalIndex` disappears as a separate method (it was the per-field wrapper), its
two rungs inlined into the loop where the index can be carried across iterations.
Verdicts are unchanged by construction — same rungs, same order, same `Ambiguous`
demotion — and `NominalResolutionParitySpec` re-confirms it: 151,515 synthetic
(schema, case-name list, declIdx) cells plus every real #98 fixture, zero
mismatches against the frozen pre-index oracle.

Measured with `com.sun.management.ThreadMXBean.getThreadAllocatedBytes`, warm,
bytes per resolution (the harness the PR's disclosed figures came from):

  shape        n=2    n=12    n=33    n=66     per case field
  identity      24      64     152     280     4 B  (the `out` array, unchanged)
  snake before 280   1,600   4,376   8,728     132 B
  snake after  256   1,336   3,608   7,168     108 B

The saving is exactly `(arity - 1) x 24 B` — one `SchemaQuery` per case field
beyond the first — confirming the key was NOT being scalar-replaced by escape
analysis. Net: -17.9% on the transformed path at n=66. What remains is the
normalised `String` itself, ~104 B per case field, which is the disclosed trade
and needs the cursor-hash probe to close. The identity-named path is
byte-identical before and after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* fix(avro): hand out an unmodifiable view of the cached name index

`normalisedNameIndex` returned the live `java.util.HashMap` it had just built
and cached. `private[avro]` is a Scala-only fence — the method is PUBLIC in
bytecode — and the map it returns is a process-wide singleton per schema behind
a `ConcurrentHashMap`, so a single in-module (or reflective) `put` on the
returned reference silently and permanently corrupts nominal field resolution
for that schema for the life of the JVM. Every later resolution would read the
corrupted entry as if the index had been built that way: wrong slot, valid wire
bytes, no error. Nothing in the type or the name warned against it.

The map is now wrapped in `Collections.unmodifiableMap`, and the WRAPPER — not
the backing map — is what goes into the cache, so the wrap is paid once per
schema at build time rather than per lookup, and repeated lookups keep handing
back the same instance.

That instance stability is what `NominalNameIndexCacheSpec` pins with
`beTheSameAs`, and it still holds: the spec's "built once per schema and handed
back by reference" and "repeated resolution never rebuilds the index" examples
pass unchanged against the stored wrapper, as does the ambiguity example, which
reads a value out of it. No test needed weakening.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* docs(avro): the injectivity comment describes the example below it

The comment on `NominalResolutionParitySpec`'s injectivity example talked about
`a`, `_a` and `A` — an earlier fixture. None of those names appear in the
example: it builds `record{a_b, b}` and asks for case fields `ab` and `aB`.

Rewritten to say what that fixture actually shows: both case fields normalise to
the key `ab`, which names the single schema field `a_b`, so the map is
non-injective and the rung abstains for EVERY field rather than letting the first
case field keep slot 0. Comment only — the example was and is correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): arm the wall-clock cost ratios only where a clock means something

`NominalResolutionCostSpec` put three timing-ratio assertions — 2n/n < 3.0,
4n/n < 10.0, snake/identity < 15.0 — into the always-run root `sbt test`, which
means they gate the temurin@17 / @21 / @25 matrix on every push and every PR,
on shared GitHub runners. That contradicts the project's own standing rule: a
nanosecond on a shared box is +/-15-50%, only B/op and within-run ratios are
load-bearing, and timing claims belong in the JMH bench pipeline. The margins
are healthy (2n/n measures 2.09-2.24, 4n/n 4.61-5.35, snake/identity 3.30-5.08)
but 3.0 against 2.18 is a 38% excursion from turning the entire matrix red for a
reason unrelated to the diff under test.

The three assertions are now armed by `-Deo.costGate=true` and skip — visibly,
with the reason printed — without it. `quality.yml` arms them, on release tags
and on demand, in a job of its OWN rather than a step in the coverage+mutation
job: a timing flake must not abort that sweep and cost the tag its QA artifacts,
a runner doing nothing else is the only place a wall-clock ratio means anything,
and a separate job never sees scoverage-instrumented bytecode, which would make
any timing meaningless. Nothing is weakened: same sizes, same bounds, same
interleaving, run where a clock means something. It is the one gate in a
workflow that is otherwise all reports, and it is isolated accordingly.

The table above them still prints on every lane. It gates nothing, costs ~4 s,
and is how a human notices drift between release-tag runs.

Also added, since the gate now runs in a fresh JVM rather than after the table:
an explicit discarded warm-up pass over all three sizes before the first timed
round. The spec is `sequential`, so the table example did warm those paths, but
that was an ordering coincidence, not a contract.

Why not the two alternatives:

  - WIDEN the bounds and keep them in the default lane. The widening would have
    to be large enough to survive an arbitrarily contended runner, at which point
    the assertion stops separating linear from quadratic — and it leaves a clock
    on the critical path of every diff in the repo. Note also that the tightest
    bound is 2n/n at 3.0, not snake/identity at 15.0; widening only the one with
    the noisy denominator would have addressed the smaller half of the problem.

  - REPLACE with a non-timing assertion. Preferred, and not available. The
    deterministic halves of the claim are already pinned: `NominalNameIndexCacheSpec`
    pins one index instance per schema and no rebuild across repeated resolutions,
    `NominalResolutionParitySpec` pins every verdict. What is left is "how much
    work per case field", which is not observable from outside the method:
    `Schema`'s constructors are package-private so it cannot be subclassed to
    count field accesses, and an internal probe counter would put mutable state on
    the hot path this issue exists to make cheap. Allocated bytes — the repo's
    other load-bearing metric, and deterministic here — cannot see the regression
    that matters either: the quadratic scan allocated LESS than the index does
    (280 B vs 7,168 B at n=66), so a return to it would read as an improvement.

The reasoning is recorded in the spec's own scaladoc, so the next reader does not
re-run the search.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* docs(changelog): the measured per-field cost, and the unmodifiable index view

The #103 entry quoted ~128 B per case field and 8,728 B at n=66. Hoisting the
cache probe out of the per-case-field loop removed one `SchemaQuery` per field
beyond the first — exactly `(arity - 1) x 24 B`, so the key was not being
scalar-replaced — and the figures are now ~104 B per case field, 24 -> 256 B at
n=2 and 280 -> 7,168 B at n=66. Also states that the identity-named path is
byte-identical before and after, which is the invariant the hoist had to keep,
and adds the unmodifiable-view entry.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
kryptt added a commit that referenced this pull request Sep 18, 2026
…68 lines

PR #98 added an all-or-nothing nominal rung to `AvroWalk.fieldNameAt`: resolve
`.field(_.x)` by NAME when every case field maps to a distinct schema field,
otherwise abstain to position. Four clauses; only TOTALITY was pinned.
Injectivity, exact-beats-normalised, and ambiguity-implies-abstain had no
fixture that could tell a correct implementation from a broken one — the sole
ambiguity fixture (`ResolutionResidualSpec` "at3") passes for the wrong reason,
because an unrelated unmapped field forces totality-driven abstention first and
masks the ambiguity branch entirely.

That is not cosmetic. Issue #104 recorded an experimental alias rung silently
breaking a case the rung gets right while all 212 avro tests stayed green: a
MIS-TARGETED optic satisfies get-put, put-get, put-put and modify fusion
perfectly, so no optic law family can see this.

`AvroNominalDoctrineSpec` states the doctrine as a declarative oracle and checks
`fieldNameAt` against it EXHAUSTIVELY over ~66k cells: every injective field
list of length 1..3 from a 6-name pool (each name both a legal Avro field name
and a legal Scala identifier), crossed with every case-name list of length 0..3
WITH repetition, crossed with declIdx from -1 to arity+1. The oracle is written
from the doctrine, not copied from the implementation — parity against a frozen
copy of the code cannot detect a wrong doctrine, which is precisely the #104
failure mode.

Enumeration, not `Gen`: the discriminating cells are threshold cells a default
generator would essentially never produce — a duplicate resolution landing on
slot 0, a fuzzy ambiguity whose FIRST hit is at index 0 versus at index >= 1,
and `declIdx` exactly equal to the schema's field count.

Verified by applying each mutant to `AvroWalk.scala` and confirming the property
goes red: 291 (exact fast path), 301 / 304 (the injectivity scan), 169 / 171 /
173 (ambiguity => abstain), 213 (`sameFieldName`'s char inequality), 310
(`declIdx >= fields.size` at the exact boundary), and 242 / 244 / 245 / 249 (the
four disjuncts of the all-or-nothing precondition).

Deletes the subsumed `"fieldNameAt: non-record parent and out-of-range declIdx
both throw loudly"` example: it probed `declIdx = 99` against a 2-field schema,
a cell where `>=` and a `>` mutant agree, so it pinned neither boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V
kryptt added a commit that referenced this pull request Sep 18, 2026
…led, 14 fewer examples (#114)

* test(avro): nominal-resolution doctrine oracle — kills 12 mutants in 68 lines

PR #98 added an all-or-nothing nominal rung to `AvroWalk.fieldNameAt`: resolve
`.field(_.x)` by NAME when every case field maps to a distinct schema field,
otherwise abstain to position. Four clauses; only TOTALITY was pinned.
Injectivity, exact-beats-normalised, and ambiguity-implies-abstain had no
fixture that could tell a correct implementation from a broken one — the sole
ambiguity fixture (`ResolutionResidualSpec` "at3") passes for the wrong reason,
because an unrelated unmapped field forces totality-driven abstention first and
masks the ambiguity branch entirely.

That is not cosmetic. Issue #104 recorded an experimental alias rung silently
breaking a case the rung gets right while all 212 avro tests stayed green: a
MIS-TARGETED optic satisfies get-put, put-get, put-put and modify fusion
perfectly, so no optic law family can see this.

`AvroNominalDoctrineSpec` states the doctrine as a declarative oracle and checks
`fieldNameAt` against it EXHAUSTIVELY over ~66k cells: every injective field
list of length 1..3 from a 6-name pool (each name both a legal Avro field name
and a legal Scala identifier), crossed with every case-name list of length 0..3
WITH repetition, crossed with declIdx from -1 to arity+1. The oracle is written
from the doctrine, not copied from the implementation — parity against a frozen
copy of the code cannot detect a wrong doctrine, which is precisely the #104
failure mode.

Enumeration, not `Gen`: the discriminating cells are threshold cells a default
generator would essentially never produce — a duplicate resolution landing on
slot 0, a fuzzy ambiguity whose FIRST hit is at index 0 versus at index >= 1,
and `declIdx` exactly equal to the schema's field count.

Verified by applying each mutant to `AvroWalk.scala` and confirming the property
goes red: 291 (exact fast path), 301 / 304 (the injectivity scan), 169 / 171 /
173 (ambiguity => abstain), 213 (`sameFieldName`'s char inequality), 310
(`declIdx >= fields.size` at the exact boundary), and 242 / 244 / 245 / 249 (the
four disjuncts of the all-or-nothing precondition).

Deletes the subsumed `"fieldNameAt: non-record parent and out-of-range declIdx
both throw loudly"` example: it probed `declIdx = 99` against a 2-field schema,
a cell where `>=` and a `>` mutant agree, so it pinned neither boundary.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): count-preserving JSON structure faults — kills 7 mutants in 13 lines

`AvroJsoniter.readRecord`'s strict exact-cover parse has three independent
refusals — unknown key, duplicate key, and the trailing arity check — plus
end-or-comma refusals in the array, map and byte-array readers. The module had
no fixture for any of the delimiter faults, and its one duplicate-key arm was
MASKED: appending `"i":42` makes 14 keys against 13 schema fields, so the arity
check refuses the document first and the `seen` scan is never exercised. Same
masking shape as the ambiguity fixture this branch's first commit replaced.

Every new arm is deliberately COUNT-PRESERVING:

  - the duplicate REPLACES a sibling key (`"l":...` becomes a second `"i":42`),
    leaving 13 keys with `l` unset, so `seen` is the only thing that can refuse;
  - the four delimiter faults close a nested record / array / map / bytes array
    with the wrong bracket, which the outer reader then resumes past, so the
    document parses end-to-end unless the inner reader refuses.

Kills 54 (`seen(field.pos) = true`), 57 and 59 (the `(field eq null) ||
seen(...)` guard and its operator), 71 (the record's objectEndOrCommaError), and
39 / 55 / 62 (the array / map / byte-array end-or-comma refusals).

Drops the now-dominated 14-key duplicate arm. The triage's proposed "unknown
JSON key" arm is NOT added: under both surviving guard mutants an unknown key
dereferences a null `field`, the NPE is caught by `parseSlice`'s `NonFatal`
handler, and the result is byte-identical to the real refusal — an equivalent
path, not a gap.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): decoded-map keys and union diagnostic payloads — kills 6 mutants in 46 lines

Three gaps in the in-memory walker, all of the same species: the existing block
asserts the SHAPE of the answer and never its content.

Map keys. The map-walk example used a hand-built `LinkedHashMap[String, _]`, so
`asMap.get(name: String)` hits directly and the `direct == null` Utf8 retry is
dead code under test. A decoded map is keyed by `Utf8` — the shape production
code actually sees — and only the retry finds anything in it. The new arm puts
the record through the binary codec and probes every key of a 6-name pool
against map sizes 0 / 1 / 2 / 5, so present and absent probes both occur at
every size. Kills 322 (the Utf8 retry) and 338 (the PathMissing refusal, whose
removal returns `Right(null)`).

Null union branch. No fixture walked `UnionBranch("null")` into a null payload,
the one branch name for which a null focus is correct rather than a failure.
Kills 191.

Union diagnostics. `UnionResolutionFailed` carries the declared alternative
list, and every existing assertion matched it with a wildcard. Recovering that
list walks back to the nearest Field step, so it needs a union two records deep
to separate the parents cursor's `pIdx >= 0` guard from a `pIdx == 0` one (kills
183), and a path whose recovered Field step names an ARRAY rather than a union
to pin the empty-list degradation — asking a non-union for its alternatives
throws an AvroRuntimeException straight out of the walk (kills 254 and 256).

Folded into the existing per-shape composite blocks per this file's stated
consolidation convention: no new registrations.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): byte-cursor refusal identities and the strictness policy — kills 11 mutants in 83 lines

Every guard in `AvroBinaryCursor`'s path walk, when removed, lets the cursor run
on into the Avro runtime, which throws, which `locateFrom` catches and reports
as `BinaryParseFailed`. A `Left` either way — so the module's `isLeft`
assertions pin none of them, and the structured diagnostic the `.record` face
publishes degrades silently to "parse failed". These rows assert WHICH
`AvroFailure` comes back, never that one does.

Called at the `AvroBinaryCursor` seam (`private[avro]`, same package) rather
than through a prism: several rows need a path the drilling macros refuse to
build, and the strictness flag is not reachable from the public surface at all.

Kills 592 (Field step against a non-record), 500 and 597 (UnionBranch step
against a non-union, and the diagnostic's own non-union guard), 601 and 605
(the ordinal scan running off the end of the alternative list), 534 (the
`requested < 0` refusal under the LENIENT policy — the strict path refuses for a
different reason and masks it), 575 (locateElements' non-array terminal) and 506
(locateElements resolving its prefix strictly).

Two rows carry their own fixtures:

  - 542, the terminal-vs-interior union split. An interior union step must not
    anchor the returned span. Every downstream observable still round-trips
    under the mutant, because the mis-anchored span still opens on a decodable
    value; only `span.valueSchema` gives it away.
  - 440, `AvroPrism`'s byte-face read resolving its terminal union strictly. The
    existing union fixtures all mismatch INTO a decode error, which `=== None`
    cannot tell apart from a correct refusal. New `Twin` fixture: two branches
    with identical field shapes and different names, so `TwinA(100)` and
    `TwinB(100)` encode to the same bytes and a lenient read hands back a
    well-formed value of the wrong type. The lenient policy is correct for the
    graft/write path only.

The triage listed 440's witness as a null-branch payload with a trailing long
field. That does not discriminate: `decodeSlice` reads a LENGTH-BOUNDED slice,
and a null branch skips zero bytes, so the mutant hits EOF and Misses exactly
like the real path. Replaced with the `Twin` fixture above.

Also pins AvroTraversal 359 — `.each`'s prefix must resolve to an array or throw
loudly. The public macro rejects a non-array focus at compile time, so the guard
is only reachable at the internal seam, which is what a future drilling entry
point would use.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): array framing survives a write — kills 4 mutants in 32 lines

The traversal's write path has three guards nobody could see: the no-op guard
(`active.isEmpty || arity mismatch`), the canonical-splice vs re-frame split,
and the initial canonical flag the block walk threads. Mutating any of them
still produces a payload that DECODES to the right value, which is why the
Optional-law property block at the foot of this file — and every other
write-side assertion in the module — leaves all four alive. What changes is the
FRAMING: a payload that arrives blocked and leaves canonical, or arrives
multi-block and leaves single-block, is a rewrite of bytes the caller never
asked us to touch.

Two cells, both asserted at byte level, and neither is reachable from the
existing fixtures:

  - all-miss under NON-canonical framing (every element on the other union
    branch). The write must be the identity on the bytes. On canonical framing
    the mutants re-frame to something byte-identical, so only the blocked
    fixture separates them. Kills 364 and 361.
  - canonical MULTI-BLOCK framing plus a length-preserving edit. Spec-legal, and
    the only shape where an in-place splice and a whole-region re-frame differ:
    collapsing two positive-count blocks into one costs a count byte, so the
    payload shrinks while decoding identically. Every other fixture here is
    single-block, where a re-frame reproduces the input exactly. Kills 356 and
    586. Hand-assembled, since neither avro encoder emits it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* test(avro): fold 27 examples into 9 — 18 fewer, 0 kills lost

The mandatory consolidation phase. Seven folds, chosen by MEASURED redundancy
rather than by reading: stryker4s records the killing test in `statusReason`,
so the report itself says which examples carry kills. Of avro's 210 registered
tests only 65 are ever the credited killer of a mutant; every example folded
below is in the other 145, which makes zero kill regression a property of the
report and not a hope — the credited killer of each of the 431 killed mutants
survives every fold untouched. It was then measured anyway.

  1. AvroCodecDecoderReuseSpec — SEVEN properties differing only in the
     (generator, writer schema, reader schema) triple and in which decode entry
     point they called, collapsed to ONE. The triple is now data
     (`reuseScenarios`), and each scenario is put through BOTH entry points —
     the thread-local cache and the `threadLocalStorage = false` opt-out —
     where before each was checked by one example each. Strictly more coverage
     per shape.
  2. AvroFieldNamingSpec — nine examples to three. The four byte-face examples
     were one fixture on one code path (fixture guard, `.field` read, `.field`
     modify, `selectDynamic`); record face and nested descent were a second
     pair. The two `.fieldNamed` examples are duplicates of
     `AvroNominalResolutionSpec`'s, which asserts strictly more of the refusal
     message, so they go. `.fields` stays its own example: it is the one shape
     in the module where the FOCUS codec's names diverge from the parent's.
  3. ConfluentReaderSpec — the two unframed-refusal examples (`reader` /
     `recordReader`) become one over a shared predicate, keeping both payload
     shapes; the MOVED-field and PROMOTED-field examples become one over a
     shared `readFramed` rig.
  4. AvroWriteCorrectnessSpec — the three `.each.union[Branch]` examples
     (blocked framing, canonical framing, record face) were one fixture and one
     optic handed to three carriers: now one example, every assertion kept.
  5. AvroWriteCorrectnessSpec — the two `.each.fields` examples likewise
     differed only in how the same basket was framed.
  6. AvroJsonSpec — the enum/bytes/fixed RENDER example and the enum/bytes/fixed
     ROUND-TRIP example built two near-identical three-leaf schemas. One schema,
     one value, and the render half now asserts the WHOLE document
     (`avroToJson(record) === json`) instead of three field lookups — stronger
     than what it replaced.

MEASURED, not assumed. `project avroIntegration; stryker` re-run and diffed
against the pre-consolidation report by the full (file, line, column, mutator,
replacement) key: 431 Killed / 38 Survived / 17 NoCoverage / 6 CompileError /
122 Ignored BEFORE, and exactly the same after, with zero keys added, zero
removed, zero Killed -> anything. Mutation score unchanged at 88.68% total /
91.90% of covered. Ran twice during the phase (after folds 1-3 and after 4-6);
both diffs clean. Against this run's original baseline the whole branch is
+40 kills with zero regressions.

No assertion was weakened and no fixture dropped: every fold keeps every
`must`/`===` it merged, and three of them add one.

avro suite: 177 -> 159 `>>` registrations, 5,201 -> 5,122 test LOC. Against the
run's starting point (173 / 4,952) that is 14 FEWER examples for +40 mutant
kills; the +170 lines are the net of 242 lines of new high-leverage tests minus
79 removed here, i.e. 4.25 net lines per additional kill.

Gates on JDK 25: `avroIntegration/test` Passed 198, Failed 0, Errors 0;
`scalafixAll; scalafmtAll` then `scalafmtCheckAll; scalafixAll --check` clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

* docs(avro): re-aim the doctrine spec's covers-block at post-#107 AvroWalk

#107 landed while this branch was open and rewrote the code the spec targets:
`sameFieldName` and the fuzzy per-field scan are gone, replaced by
`normalisedName` plus a weak-keyed `normalisedNameIndex` cache. The spec itself
needed no change — it was written against the public seam `AvroWalk.fieldNameAt`
precisely so it would survive an internals rewrite, and it agrees with the new
implementation cell for cell — but its `covers:` block named deleted symbols and
pre-#107 line numbers.

Re-aims it at the surviving constructs by name and drops the line numbers, and
states the division of labour with #107's own `NominalResolutionOracle`: that
one is the OLD ALGORITHM frozen verbatim (a refactor gate — "the new rung agrees
with the one it replaced"), this one is written from the DOCTRINE without
reading the implementation. Parity cannot detect a wrong doctrine, which is the
#104 failure mode; the doctrine oracle still passing against wholly rewritten
internals is the evidence that #107 preserved the contract and not just the code
path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0194EHFR4NamCpTHiqy7B74V

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant