Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
55 changes: 55 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,61 @@ All notable changes to this project will be documented in this file.
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/),
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).

## [Unreleased]

### Fixed

- **`.field(_.x)` no longer targets the wrong schema field when the codec's schema is not
positionally 1:1 with the case class** (#95): resolution was `fields.get(declIdx).name` and
nothing else — the field's NAME, its TYPE and the record's ARITY were never consulted. Sound for
every kindlings-derived codec (1:1 by construction), unsound for a hand-written or `vulcan.Codec`
field list, which can add a computed column, drop one, or reorder. On those, `.field(_.b)` against
a schema `{a, computed, b, c}` read and wrote `computed`, producing **valid Avro bytes with wrong
content** — and no law could see it, because a mis-targeted optic is a perfectly lawful `Optional`
onto the wrong field. The same resolver backs `.fields`, `selectDynamic`, `.each.field` and
`.each.fields`, so all six sites were affected. Resolution now tries a NAME rung first, and only
when the codec has named the WHOLE case-field list onto distinct schema fields (total and
injective, exact or up to `_`/`-`/`.` and case); otherwise it abstains and declaration position
decides exactly as before — which is what keeps every name-transform codec (issue #35's
population) resolving correctly. Still construction-time only: zero per-operation cost.
- **`.fieldNamed("typo")` is refused at construction instead of silently missing at runtime**
(#95): the explicit-schema-name escape hatch appended the literal with no schema lookup at all,
although the schema was in hand, so a typo — or the Scala field name passed where the schema name
was meant — read `None` and wrote the payload back unchanged while reporting success. That is the
same failure class the hatch exists to avoid, and it is what every "navigate by explicit schema
name with `.fieldNamed`" error message points at. Map parents are carved out: `.fieldNamed` is
also how a map KEY is addressed, and an absent key is data. Feature-detecting a RECORD field is
still available via `codec.schema.getField(name)`.
- **A nested selector `.field(_.a.b)` is now a compile error** (#95): the selector parser shared by
every cursor macro matched `Lambda(_, Select(_, name))` with ANY receiver, so `_.inner.y` parsed
as the bare name `y` and the macro resolved it on the PARENT. Where the parent carries a field of
that name — and a record holding a nested record often does — the result was a well-typed,
perfectly lawful optic aimed at the wrong field: silent corruption on a 1:1, derived codec, with
no schema divergence involved, and the macros' own "nested paths … chain them" abort unreachable
for exactly the shape it was written for. A single-hop selector naming something that is not a
case field of the parent (a no-arg `def`) is now a compile error too, instead of a literal field
name that misses at runtime. `AvroPrism` / `AvroTraversal`, `JsonPrism` / `JsonTraversal`
(eo-circe) and `JsoniterPrism` / `JsoniterTraversal` (eo-jsoniter) all share the parser, so all
six surfaces are covered.

### Changed

- **Behaviour change, avro**: a hand-written codec that PERMUTES the Scala names (writes case field
`a` into a schema field literally named `b`, and vice versa) resolved correctly by position and
now resolves by name, i.e. wrongly. No name transform can produce that shape — a transform is a
function of the name alone — but a hand-written field list can. Use `.fieldNamed` there.
- **Recompile, do not re-jar**: the avro resolution signatures are `private[avro]`, but
`transparent inline` bakes the accessor into CALLER bytecode, so downstream projects must
recompile against this release rather than swapping the jar.

### Known limitations

Three codec shapes are still resolved to the wrong schema field, unchanged from 0.15.1 and pinned
as executable examples in `ResolutionResidualSpec`: a field list that both renames beyond
recognition and reorders; a schema column that bears a case field's name but holds a different
value (a derived public id, a stale legacy column); and two columns whose names normalise alike.
`.fieldNamed("schema_name")` reaches all of them.

## [0.15.1] - 2026-08-20

### Fixed
Expand Down
97 changes: 70 additions & 27 deletions avro/src/main/scala/dev/constructive/eo/avro/AvroPrism.scala
Original file line number Diff line number Diff line change
Expand Up @@ -42,15 +42,38 @@ import org.apache.avro.Schema
* drilled-prism given is the evidence that instantiates their `T` at `Array[Byte]`. See the
* docs page's migration recipe for the runnable shape.
*
* '''Field navigation honours the SCHEMA field name (issue #35).''' `.field(_.x)` (and `.fields`,
* `selectDynamic`, the traversal siblings) resolve the case-class field `x` to whatever schema
* field the codec actually emitted for it — under any name transform (kindlings snake / kebab /
* custom `transformFieldNames`, or a vulcan per-field override map) — by DECLARATION POSITION: the
* i-th case field maps to the i-th schema field, read back out of the cached schema at
* construction time (zero per-operation cost). The rare hand-written codec whose schema field
* ORDER diverges from declaration order needs [[AvroPrism.fieldNamed]]`("schema_name")` to
* navigate by the explicit schema name instead. Map keys are data, not schema-named fields, and
* keep their literal key.
* '''Field navigation honours the SCHEMA field name (issues #35, #95).''' `.field(_.x)` — and
* equally `.fields(...)`, `selectDynamic` and the `.each.field` / `.each.fields` traversal
* siblings, which all share one resolver — maps the case-class field `x` to whatever schema field
* the codec actually emitted for it. Resolution happens ONCE, at prism construction, off the
* cached schema (zero per-operation cost), by two rungs:
*
* 1. '''By NAME, all-or-nothing.''' If EVERY case field of the parent maps to a DISTINCT schema
* field — exactly, or uniquely up to `_` / `-` / `.` and case — then the codec has named the
* whole correspondence and `x`'s answer is read off that map. Partial or colliding coverage
* is treated as no signal at all and the rung abstains for every field, because one lucky
* name match on a schema whose OTHER columns are legacy is how a working call site gets
* re-aimed at the wrong column.
* 1. '''By DECLARATION POSITION''' — the i-th case field is the i-th schema field. This is where
* a name transform lands (a kindlings snake-case config, a custom `transformFieldNames`, a
* vulcan per-field override map), because a transform REMOVES the literal Scala name by
* construction, so rung 1 cannot have fired.
*
* '''The positional rung is only right when the codec's schema is positionally 1:1 with the case
* class.''' Kindlings-derived codecs are, by construction. A hand-written or `vulcan.Codec` field
* list need not be: a COMPUTED/derived schema column, a dropped field, or a reordered field list
* all break it. The name rung recovers most of that population; what it cannot recover is a field
* list that both renames beyond recognition AND reorders (equal arity, no name hit) — that
* resolves by position, silently, and is wrong. Two more shapes stay wrong for the same reason: a
* schema column that BEARS a case field's name but HOLDS a different value (a derived public id, a
* stale legacy column), and two columns whose names normalise alike. Navigate all of them with
* [[AvroPrism.fieldNamed]]`("schema_name")`, which bypasses resolution entirely and is itself
* checked against the schema; `ResolutionResidualSpec` pins each shape's exact behaviour.
*
* Behaviour change against 0.15.1, for the release notes: a hand-written codec that PERMUTES the
* Scala names (writes case field `a` into a schema field literally named `b`, and vice versa)
* resolved correctly by position and now resolves by name, i.e. wrongly. No name transform can
* produce that shape. Map keys are data, not schema-named fields, and keep their literal key.
*
* Two sibling surfaces, one mechanism each (deliberately NOT duplicated here):
*
Expand All @@ -75,14 +98,16 @@ import org.apache.avro.Schema
* `.replace` onto a span whose current value doesn't decode as `A` — or a `.union[B]` focus
* sitting on a different runtime branch — is a Miss pass-through. [[graftBytes]] is the
* decode-free write (and the only one that can SWITCH union branches).
* - '''The payload must be encoded under exactly this prism's reader schema.''' The byte walk
* performs no writer/reader schema resolution: structurally drifted payloads Miss silently,
* and a same-typed field REORDER between writer and reader is undetectable from the bytes —
* the walk reads the wrong field with full confidence. Confluent-framed payloads are handled
* by composing [[ConfluentWire.confluent]] (a byte Prism that strips the header, resolves the
* writer schema, and fingerprint-gates) BEFORE this optic — `confluent.andThen(thisWalk)`;
* past a fingerprint mismatch a mixed-schema topic still needs a resolving decode (the record
* face with the right schema per payload).
* - '''The payload must be encoded under exactly this prism's reader schema.''' This is about
* PAYLOAD drift — a name absent from the READER schema is a construction-time refusal on both
* `.field` and `.fieldNamed`, not a runtime miss. The byte walk performs no writer/reader
* schema resolution: structurally drifted payloads Miss silently, and a same-typed field
* REORDER between writer and reader is undetectable from the bytes — the walk reads the wrong
* field with full confidence. Confluent-framed payloads are handled by composing
* [[ConfluentWire.confluent]] (a byte Prism that strips the header, resolves the writer
* schema, and fingerprint-gates) BEFORE this optic — `confluent.andThen(thisWalk)`; past a
* fingerprint mismatch a mixed-schema topic still needs a resolving decode (the record face
* with the right schema per payload).
* - Dynamic field sugar is shadowed by real members: an Avro field named like a member of this
* class (`record`, `field`, `at`, `union`, `each`, `fields`, …) must be drilled with the
* explicit `.field(_.record)` form.
Expand Down Expand Up @@ -241,16 +266,24 @@ final class AvroPrism[A] private[avro] (
// ---- Path widening (used by macro extensions) ---------------------

/** Extend the Leaf path by a field step. Used by [[field]] / `selectDynamic`. `scalaName` is the
* case-class field name and `declIdx` its declaration index; the actual schema field name (which
* may differ under a snake/kebab/custom transform or vulcan overrides) is resolved off the
* cached schema by position — see [[AvroWalk.resolveFieldName]] (issue #35).
* case-class field name, `declIdx` its declaration index and `caseNames` the parent's whole
* case-field list; the actual schema field name (which may differ under a snake/custom transform
* or vulcan overrides) is resolved off the cached schema by the name-then-position rule — see
* [[AvroWalk.fieldNameAt]] (issues #35 and #95).
*/
private[avro] def widenPath[B](scalaName: String, declIdx: Int)(using
private[avro] def widenPath[B](scalaName: String, declIdx: Int, caseNames: List[String])(using
codecB: AvroCodec[B]
): AvroPrism[B] =
widenPathStep[B](
PathStep.Field(
AvroWalk.resolveFieldName(rootSchemaCached, path, scalaName, declIdx, "AvroPrism.field")
AvroWalk.resolveFieldName(
rootSchemaCached,
path,
scalaName,
declIdx,
caseNames,
"AvroPrism.field",
)
)
)

Expand All @@ -261,6 +294,7 @@ final class AvroPrism[A] private[avro] (
private[avro] def widenPathNamed[B](schemaName: String)(using
codecB: AvroCodec[B]
): AvroPrism[B] =
AvroWalk.requireFieldNamed(rootSchemaCached, path, schemaName, "AvroPrism.fieldNamed")
widenPathStep[B](PathStep.Field(schemaName))

/** Extend by an array-index step. Used by [[at]]. */
Expand Down Expand Up @@ -289,9 +323,10 @@ final class AvroPrism[A] private[avro] (
private[avro] def toFieldsPrism[B](
scalaNames: Array[String],
declIdxs: Array[Int],
caseNames: List[String],
)(using codecB: AvroCodec[B]): AvroPrism[B] =
new AvroPrism[B](
new AvroFocus.Fields[B](path, resolveFieldNames(scalaNames, declIdxs), codecB),
new AvroFocus.Fields[B](path, resolveFieldNames(scalaNames, declIdxs, caseNames), codecB),
rootSchemaCached,
)

Expand All @@ -301,13 +336,15 @@ final class AvroPrism[A] private[avro] (
private def resolveFieldNames(
scalaNames: Array[String],
declIdxs: Array[Int],
caseNames: List[String],
): Array[String] =
Array.tabulate(scalaNames.length)(i =>
AvroWalk.resolveFieldName(
rootSchemaCached,
path,
scalaNames(i),
declIdxs(i),
caseNames,
"AvroPrism.fields",
)
)
Expand Down Expand Up @@ -339,10 +376,16 @@ object AvroPrism:
)(using codecB: AvroCodec[B]): AvroPrism[B] =
${ AvroPrismMacro.fieldImpl[A, B]('o, 'selector, 'codecB) }

/** `.fieldNamed[B]("schema_name")` — drill by the EXPLICIT schema field name, bypassing position
* resolution. The escape hatch (issue #35) for a hand-written codec whose schema field order
* diverges from case-class declaration order; the common (derived / order-preserving) codecs
* need `.field(_.x)` instead, which resolves the name for you.
/** `.fieldNamed[B]("schema_name")` — drill by the EXPLICIT schema field name, bypassing
* resolution entirely. The escape hatch for a hand-written codec the resolver cannot read (a
* field list that both renames beyond recognition and reorders, a column bearing another field's
* name, two columns normalising alike); the common (derived / order-preserving /
* name-transformed) codecs need `.field(_.x)` instead, which resolves the name for you.
*
* The name is CHECKED against the schema it will be looked up in, at construction (issue #95): a
* name the record does not carry throws rather than Missing silently at runtime. A MAP parent is
* carved out — `.fieldNamed` is also how a map KEY is addressed, and an absent key is data, not
* a mistake. To feature-detect a RECORD field, ask the schema: `codec.schema.getField(name)`.
*/
extension [A](o: AvroPrism[A])

Expand Down
61 changes: 54 additions & 7 deletions avro/src/main/scala/dev/constructive/eo/avro/AvroPrismMacro.scala
Original file line number Diff line number Diff line change
Expand Up @@ -25,13 +25,18 @@ object AvroPrismMacro:
report.errorAndAbort(
"AvroPrism.field: selector must be a single-field accessor like `_.fieldName`.\n"
+ "Nested paths are not yet supported inside a single call;\n"
+ "chain them: `_.field(_.a).field(_.b)`.\n"
+ "chain them: `.field(_.a).field(_.b)`.\n"
+ s"Got: ${selector.asTerm.show}"
)
}
MacroSelectors.requireCaseField[A]("AvroPrism.field", name)

'{
$parent.widenPath[B](${ Expr(name) }, ${ Expr(declIndexOf[A](name)) })(using $codecB)
$parent.widenPath[B](
${ Expr(name) },
${ Expr(declIndexOf[A](name)) },
${ Expr(caseNamesOf[A]) },
)(using $codecB)
}

/** Drives `codecPrism[Person].name`. Looks `name` up on `A`'s schema, summons `AvroCodec[B]`,
Expand All @@ -45,7 +50,13 @@ object AvroPrismMacro:
"AvroPrism selectDynamic",
nameE,
) { [b] => (name: String, declIdx: Int, codecB: Expr[AvroCodec[b]]) =>
'{ $parent.widenPath[b](${ Expr(name) }, ${ Expr(declIdx) })(using $codecB) }
'{
$parent.widenPath[b](
${ Expr(name) },
${ Expr(declIdx) },
${ Expr(caseNamesOf[A]) },
)(using $codecB)
}
}

/** Macro for `.at(i)`. Verifies `A <: Iterable`, extracts the element type, summons the codec,
Expand Down Expand Up @@ -199,12 +210,19 @@ object AvroPrismMacro:
val name: String = MacroSelectors.extractFieldName(selector.asTerm).getOrElse {
report.errorAndAbort(
"AvroTraversal.field: selector must be a single-field accessor like `_.fieldName`.\n"
+ "Nested paths are not yet supported inside a single call;\n"
+ "chain them: `.field(_.a).field(_.b)`.\n"
+ s"Got: ${selector.asTerm.show}"
)
}
MacroSelectors.requireCaseField[A]("AvroTraversal.field", name)

'{
$parent.widenSuffix[B](${ Expr(name) }, ${ Expr(declIndexOf[A](name)) })(using $codecB)
$parent.widenSuffix[B](
${ Expr(name) },
${ Expr(declIndexOf[A](name)) },
${ Expr(caseNamesOf[A]) },
)(using $codecB)
}

/** Traversal counterpart to [[atImpl]] — extends the suffix by an array index. */
Expand Down Expand Up @@ -235,7 +253,14 @@ object AvroPrismMacro:
namesExpr: Expr[Array[String]],
declIdxsExpr: Expr[Array[Int]],
codecNT: Expr[AvroCodec[nt]],
) => '{ $parent.toFieldsPrism[nt]($namesExpr, $declIdxsExpr)(using $codecNT) }
) =>
'{
$parent.toFieldsPrism[nt](
$namesExpr,
$declIdxsExpr,
${ Expr(caseNamesOf[A]) },
)(using $codecNT)
}
}

/** Traversal counterpart to [[fieldsImpl]]. */
Expand All @@ -248,7 +273,14 @@ object AvroPrismMacro:
namesExpr: Expr[Array[String]],
declIdxsExpr: Expr[Array[Int]],
codecNT: Expr[AvroCodec[nt]],
) => '{ $parent.toFieldsTraversal[nt]($namesExpr, $declIdxsExpr)(using $codecNT) }
) =>
'{
$parent.toFieldsTraversal[nt](
$namesExpr,
$declIdxsExpr,
${ Expr(caseNamesOf[A]) },
)(using $codecNT)
}
}

/** Traversal counterpart to [[selectFieldImpl]] — drives Dynamic sugar by extending the suffix.
Expand All @@ -261,7 +293,13 @@ object AvroPrismMacro:
"AvroTraversal selectDynamic",
nameE,
) { [b] => (name: String, declIdx: Int, codecB: Expr[AvroCodec[b]]) =>
'{ $parent.widenSuffix[b](${ Expr(name) }, ${ Expr(declIdx) })(using $codecB) }
'{
$parent.widenSuffix[b](
${ Expr(name) },
${ Expr(declIdx) },
${ Expr(caseNamesOf[A]) },
)(using $codecB)
}
}

/** Shared backbone for [[fieldsImpl]] / [[fieldsTraversalImpl]] — validation + SELECTOR-order
Expand Down Expand Up @@ -318,6 +356,15 @@ object AvroPrismMacro:
import quotes.reflect.*
TypeRepr.of[A].typeSymbol.caseFields.indexWhere(_.name == name)

/** `A`'s case-field names in declaration order (`Nil` when `A` isn't a case class — a NamedTuple
* parent, say). Emitted as a compile-time literal list beside the declaration index so
* construction-time resolution can check whether the codec's schema names the WHOLE case-field
* list before it trusts any one of them (issue #95). `Nil` abstains, leaving position in charge.
*/
private def caseNamesOf[A: Type](using q: Quotes): List[String] =
import quotes.reflect.*
TypeRepr.of[A].typeSymbol.caseFields.map(_.name)

/** Summon `AvroCodec[B]` with a caller-supplied error message. */
private def summonCodec[B: Type](
errorMsg: String
Expand Down
Loading