From a9b9c2b8983b0b11cdb6caca0653570c5e7ceb61 Mon Sep 17 00:00:00 2001 From: Rodolfo Hansen Date: Tue, 29 Sep 2026 00:50:34 +0200 Subject: [PATCH] =?UTF-8?q?perf(avro):=20reuse=20write=20plumbing=20in=20w?= =?UTF-8?q?riteDatum=20=E2=80=94=20attribution=20+=20fix=20for=20#119?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Issue #119 measured the kindlings-derived AvroCodec encode route at 1.4-1.7x a hand-written direct-BinaryEncoder writer and 9-18% more allocated. An attribution decomposition (docs/research note) shows the time gap is tree build + GenericDatumWriter dispatch (kindlings' side), but most of the ALLOCATION gap was ours: writeDatum allocated a fresh ByteArrayOutputStream (32-byte start), GenericDatumWriter, and BufferedBinaryEncoder on every call — 2,928 B/op of churn on a 245 B record. writeDatum now writes through a per-thread DatumWriters (reused BAOS + re-bound BufferedBinaryEncoder + per-schema cached writers), mirroring the read side's DatumReaders/binaryDecoderCache. Byte output is unchanged; AvroWriteCorrectnessSpec pins retention under same-thread churn, recovery after an aborted write, and 8-thread isolation. AvroEncodeRouteBench (mapped into the bench CI) keeps the four routes permanently comparable; measured write-side cost drops to just the returned array (full route 3,456 -> 904 B/op). Also checked and refuted, in the research note, the bundled claim that generics lens[S] reorders multi-selector focus types — selector order is pinned by type ascriptions against published 0.14.0 and 0.17.0. --- .github/bench/bench_tools.py | 1 + .../eo/avro/AvroBinaryCursor.scala | 72 +++++- .../eo/avro/AvroWriteCorrectnessSpec.scala | 52 ++++ .../eo/bench/AvroEncodeRouteBench.scala | 222 ++++++++++++++++++ ...09-29-issue-119-avro-encode-attribution.md | 113 +++++++++ 5 files changed, 451 insertions(+), 9 deletions(-) create mode 100644 benchmarks/src/main/scala/dev/constructive/eo/bench/AvroEncodeRouteBench.scala create mode 100644 docs/research/2026-09-29-issue-119-avro-encode-attribution.md diff --git a/.github/bench/bench_tools.py b/.github/bench/bench_tools.py index a0778a21..ecefaf40 100644 --- a/.github/bench/bench_tools.py +++ b/.github/bench/bench_tools.py @@ -42,6 +42,7 @@ MODULE_BENCHES = { "avro/": [ "AvroBytesBench", + "AvroEncodeRouteBench", "AvroJsonBridgeBench", "AvroVulcanBench", "OrderAvroBench", diff --git a/avro/src/main/scala/dev/constructive/eo/avro/AvroBinaryCursor.scala b/avro/src/main/scala/dev/constructive/eo/avro/AvroBinaryCursor.scala index b251352c..492aff09 100644 --- a/avro/src/main/scala/dev/constructive/eo/avro/AvroBinaryCursor.scala +++ b/avro/src/main/scala/dev/constructive/eo/avro/AvroBinaryCursor.scala @@ -7,7 +7,14 @@ import java.io.{ByteArrayOutputStream, InputStream} import java.util.{Arrays, HashMap, List as JList} import org.apache.avro.Schema import org.apache.avro.generic.{GenericDatumReader, GenericDatumWriter, IndexedRecord} -import org.apache.avro.io.{BinaryData, BinaryDecoder, Decoder, DecoderFactory, EncoderFactory} +import org.apache.avro.io.{ + BinaryData, + BinaryDecoder, + BinaryEncoder, + Decoder, + DecoderFactory, + EncoderFactory +} /** Internal byte-offset locator behind [[AvroPrism]]'s byte-carried optic (`to`/`from`) and its * slice/graft surface. @@ -357,17 +364,64 @@ private[avro] object AvroBinaryCursor: */ private[avro] val leaves = new DatumReaders[Any] + /** Per-thread reusable binary-write plumbing — [[writeDatum]]'s engine, the write mirror of the + * [[DatumReaders]] reader cache + [[binaryDecoderCache]] decoder on the read side (issue #119). + * + * '''Why cache at all:''' every whole-record and leaf write funnels through [[writeDatum]]. Its + * first form allocated a fresh `ByteArrayOutputStream` (32-byte start — a multi-KB payload + * reallocates and copies its buffer a dozen times on the way up), a fresh `GenericDatumWriter`, + * and a fresh `BufferedBinaryEncoder` (2 KB internal buffer) on EVERY call. Measured on a + * 15-field record with an all-`Option` nested record and an 18-branch union (245 B output): + * 2,928 B/op of write-side plumbing against a 392 B/op reused-plumbing floor and 397 B/op for a + * hand-written direct-to-encoder writer — the plumbing, not `GenericDatumWriter` dispatch, is + * the bulk of the allocation gap reported in issue #119 over hand-written producers. + * + * '''Rebind, not rebuild:''' `EncoderFactory.binaryEncoder(out, reuse)` reconfigures a + * `BufferedBinaryEncoder` in place — position reset to 0, same 2 KB buffer (only replaced when + * the factory buffer size changes) — so steady-state writes allocate exactly the returned + * `toByteArray` result. The writer cache is keyed per `Schema` like the read cache is keyed per + * schema pair; `GenericDatumWriter` is mutable and NOT thread-safe, hence the `ThreadLocal` — + * the same reason [[DatumReaders.cache]] exists. + * + * '''Lifetime caveat (same as [[DatumReaders]]):''' the writer map has no eviction — it grows by + * one entry per distinct schema written on the thread. Producers write one schema repeatedly, so + * this is the intended shape; dynamically-built schemas on a huge pool would instead want fresh + * writers per call, which is exactly what the pre-#119 form did. + */ + final private class DatumWriters: + + private val out = new ByteArrayOutputStream(4096) + private var encoder: BinaryEncoder = EncoderFactory.get().binaryEncoder(out, null) + private val cache = new HashMap[Schema, GenericDatumWriter[Any]]() + + /** Encode `datum` (of `schema` shape) to a fresh, detached `Array[Byte]`. Safe against a + * mid-write throw by ordering, not by state: `out` is reset at the START of every encode and + * the previous result was already copied out by `toByteArray`, so dirty bytes from an aborted + * write are never observable. + */ + def write(datum: Any, schema: Schema): Array[Byte] = + out.reset() + encoder = EncoderFactory.get().binaryEncoder(out, encoder) + cache + .computeIfAbsent(schema, s => new GenericDatumWriter[Any](s)) + .write(datum, encoder) + encoder.flush() + out.toByteArray + + end DatumWriters + + /** One [[DatumWriters]] plumbing set per writing thread — see its scaladoc for why reuse is + * thread-local. + */ + private val writeCache: ThreadLocal[DatumWriters] = + ThreadLocal.withInitial(() => new DatumWriters) + /** THE module's binary write — [[DatumReaders.read]]'s mirror: encode an `Any`-shaped `datum` - * under `schema` to its binary wire form. Fresh writer/encoder per call: `GenericDatumWriter` - * carries no resolution state worth caching. + * under `schema` to its binary wire form. Delegates to the per-thread [[DatumWriters]] plumbing + * (issue #119); the returned array is always freshly copied. */ private[avro] def writeDatum(datum: Any, schema: Schema): Array[Byte] = - val out = new ByteArrayOutputStream() - val writer = new GenericDatumWriter[Any](schema) - val encoder = EncoderFactory.get().binaryEncoder(out, null) - writer.write(datum, encoder) - encoder.flush() - out.toByteArray + writeCache.get().write(datum, schema) /** Read a `ByteBuffer`'s remaining bytes without disturbing its position — how a `bytes` field * arrives in the generic runtime model. Shared by the bridges' structural walks (`AvroJson` / diff --git a/avro/src/test/scala/dev/constructive/eo/avro/AvroWriteCorrectnessSpec.scala b/avro/src/test/scala/dev/constructive/eo/avro/AvroWriteCorrectnessSpec.scala index 6f57f6c5..7e13c9d0 100644 --- a/avro/src/test/scala/dev/constructive/eo/avro/AvroWriteCorrectnessSpec.scala +++ b/avro/src/test/scala/dev/constructive/eo/avro/AvroWriteCorrectnessSpec.scala @@ -479,4 +479,56 @@ class AvroWriteCorrectnessSpec extends Specification with ScalaCheck: codecPrism[FullName].field(_.first).getOption(bytes) === Some("Doe") } + // ---- reused write plumbing (issue #119) ----------------------------- + + // covers: writeDatum's per-thread reused ByteArrayOutputStream / BinaryEncoder / + // GenericDatumWriter is safe on the three axes reuse introduces: + // (a) RETENTION — a result array handed to the caller is never disturbed by later writes on + // the same thread (every result is copied out by toByteArray); note the byte-face optic + // contracts already promise freshly-returned arrays. + // (b) ABORTED WRITE — a datum/schema mismatch throws mid-encode (surfaced as Left); the + // dirty buffer it leaves behind must never leak into the next write's bytes. + // (c) THREAD ISOLATION — the plumbing is ThreadLocal, so concurrent writers interleave + // freely and every thread's bytes equal the single-writer golden bytes. + "reused write plumbing: retention, aborted write, thread isolation (issue #119)" >> { + val pc = summon[AvroCodec[Person]] + val p = Person("Alice", 42) + val q = Person("Bob", 404) + + def encode(x: Person): Array[Byte] = + AvroCodec.encodeValue(x)(using pc).getOrElse(throw new RuntimeException("encode failed")) + + val goldenP = encode(p) + val goldenQ = encode(q) + val snapshotP = goldenP.clone() + val snapshotQ = goldenQ.clone() + + // 200 interleaved writes on THIS thread must not disturb the two retained results. + (1 to 200).foreach { n => encode(if n % 2 == 0 then p else q); () } + val churnOk = Arrays.equals(goldenP, snapshotP) && Arrays.equals(goldenQ, snapshotQ) + + val aborted = AvroCodec.encodeRecord("definitely not a record", pc.schema) + val cleanAfterAbort = Arrays.equals(goldenP, encode(p)) + + val results = new java.util.concurrent.ConcurrentLinkedQueue[Boolean]() + val writers = (1 to 8).toList.map { _ => + val t = new Thread(() => + (1 to 50).foreach { n => + val x = if n % 2 == 0 then p else q + val golden = if n % 2 == 0 then goldenP else goldenQ + results.add(Arrays.equals(golden, encode(x))) + } + ) + t.start() + t + } + writers.foreach(_.join()) + + (aborted.isLeft === true) + .and(churnOk === true) + .and(cleanAfterAbort === true) + .and(results.size() === 400) + .and(results.contains(false) === false) + } + end AvroWriteCorrectnessSpec diff --git a/benchmarks/src/main/scala/dev/constructive/eo/bench/AvroEncodeRouteBench.scala b/benchmarks/src/main/scala/dev/constructive/eo/bench/AvroEncodeRouteBench.scala new file mode 100644 index 00000000..184235fd --- /dev/null +++ b/benchmarks/src/main/scala/dev/constructive/eo/bench/AvroEncodeRouteBench.scala @@ -0,0 +1,222 @@ +package dev.constructive.eo +package bench + +import scala.compiletime.uninitialized + +import java.io.ByteArrayOutputStream +import java.util.concurrent.TimeUnit + +import avro.AvroCodec +import hearth.kindlings.avroderivation.{AvroConfig, AvroDecoder, AvroEncoder, AvroSchemaFor} +import org.apache.avro.Schema +import org.apache.avro.generic.GenericDatumWriter +import org.apache.avro.io.{BinaryEncoder, EncoderFactory} +import org.openjdk.jmh.annotations.* +import scala.jdk.CollectionConverters.* + +/** Fixtures + routes for [[AvroEncodeRouteBench]]. Kept top-level: kindlings' derivation, like + * hearth's constructor synthesis, must not see an outer accessor. + */ +object EncodeRouteImpls: + + given AvroConfig = AvroConfig() + + /** All-`Option` nested record — the issue's 55-field `metrics` shape, scaled to 6. */ + final case class Metrics( + a1: Option[Double], + a2: Option[Double], + b1: Option[Long], + s1: Option[String], + i1: Option[Int], + b2: Option[Boolean], + ) + + object Metrics: + given AvroEncoder[Metrics] = AvroEncoder.derived + given AvroDecoder[Metrics] = AvroDecoder.derived + given AvroSchemaFor[Metrics] = AvroSchemaFor.derived + + /** Multi-branch union — the issue's 18-branch slot, scaled to 6. */ + enum Event: + case Ev0(v: Long) + case Ev1(v: Long) + case Ev2(v: Long) + case Ev3(v: Long) + case Ev4(v: Long) + case Ev5(v: Long) + + object Event: + given AvroEncoder[Event] = AvroEncoder.derived + given AvroDecoder[Event] = AvroDecoder.derived + given AvroSchemaFor[Event] = AvroSchemaFor.derived + + /** 10-field top level: strings, numerics, a boolean, one optional nested record, one union. */ + final case class Payload( + id: String, + tenant: String, + source: String, + ts: Long, + seq: Long, + amount: Double, + flag: Boolean, + kind: Int, + metrics: Option[Metrics], + event: Event, + ) + + object Payload: + given AvroEncoder[Payload] = AvroEncoder.derived + given AvroDecoder[Payload] = AvroDecoder.derived + given AvroSchemaFor[Payload] = AvroSchemaFor.derived + + val payload: Payload = + Payload( + "id-1234567890", + "tenant-xyz", + "src.system.a", + 1_700_000_000L, + 987_654L, + 123_456.789, + flag = true, + kind = 3, + metrics = Some(Metrics(Some(1.5), Some(2.5), None, Some("alpha"), None, Some(true))), + event = Event.Ev3(99L), + ) + + val codec: AvroCodec[Payload] = summon[AvroCodec[Payload]] + val schema: Schema = codec.schema + val metricsUnion: Schema = schema.getField("metrics").schema() + + /** Which arm of a `union` is the null arm — read off the DERIVED schema once at object + * init, so the hand-written route below never guesses the spelling and never re-derives it on + * the hot path. + */ + def nullArm(s: Schema): Int = + s.getTypes.asScala.indexWhere(_.getType == Schema.Type.NULL) + + private def nonNullArm(s: Schema): Schema = + s.getTypes.asScala.find(_.getType != Schema.Type.NULL).get + + val metricsSomeIdx: Int = 1 - nullArm(metricsUnion) + private val metricsSchema: Schema = nonNullArm(metricsUnion) + + val metricsNullIdx: Array[Int] = + List("a1", "a2", "b1", "s1", "i1", "b2") + .map(n => nullArm(metricsSchema.getField(n).schema())) + .toArray + +/** Whole-record ENCODE attribution — the route decomposition issue #119 asked for. + * + * The report measured `AvroCodec.derived` encode at ~1.4–1.7× a hand-written + * direct-`BinaryEncoder` writer and ~9–18% more allocated, and hypothesised the intermediate + * `GenericData.Record` tree. Four routes split the cost so the remaining gap is attributable: + * + * - `eo_encodeValue` — the full production route: kindlings A → Any, then the module's + * `writeDatum`. Post-#119 `writeDatum` writes through per-thread reused buffer/encoder/writer + * plumbing, so its allocation over `eo_encodeToAny` is (almost) exactly the returned `byte[]`. + * - `naive_freshPlumbing` — the pre-#119 shape reproduced here as a baseline: a FRESH + * `ByteArrayOutputStream` (32-byte start, doubling), `GenericDatumWriter` and + * `BufferedBinaryEncoder` per call over the SAME prebuilt tree. The delta vs `eo_encodeValue` + * is the per-call plumbing churn the fix removed — the dominant write-side allocator, not + * `GenericDatumWriter` dispatch. + * - `eo_encodeToAny` — the kindlings tree build alone (the reporter's hypothesis: what the + * generic-record materialisation costs, isolated). + * - `handwritten_stream` — the reporter's baseline: fields straight to a reused `BinaryEncoder` + * in schema order, no tree. The gap to `eo_encodeValue` is tree + `GenericDatumWriter` + * dispatch — the irreducible cost of going through avro's datum model, and what a true + * streaming `A => Encoder => Unit` derivation would eventually remove. + * + * Run with the GC profiler — B/op is the metric, ns/op advises: + * {{{ + * sbt "benchmarks/Jmh/run -i 5 -wi 3 -f 3 -t 1 -prof gc .*AvroEncodeRouteBench.*" + * }}} + */ +@State(Scope.Benchmark) +@BenchmarkMode(Array(Mode.AverageTime)) +@OutputTimeUnit(TimeUnit.NANOSECONDS) +@Fork(3) +@Warmup(iterations = 3, time = 1) +@Measurement(iterations = 5, time = 1) +class AvroEncodeRouteBench extends JmhDefaults: + + import EncodeRouteImpls.* + + var tree: Any = uninitialized + var out: ByteArrayOutputStream = uninitialized + var encoder: BinaryEncoder = uninitialized + + @Setup(Level.Trial) + def init(): Unit = + tree = codec.encode(payload) + out = new ByteArrayOutputStream(16384) + encoder = EncoderFactory.get().binaryEncoder(out, null) + // sanity: the hand-written route must land on the codec's bytes, or the attribution is fiction + val eo = AvroCodec.encodeValue(payload)(using codec).getOrElse(null) + val hand = handwrittenToByteArray() + require( + eo != null && java.util.Arrays.equals(eo, hand), + "handwritten route diverged from eo encode" + ) + + @Benchmark def eo_encodeValue: Array[Byte] = + AvroCodec.encodeValue(payload)(using codec).fold(_ => null, identity) + + @Benchmark def eo_encodeToAny: Any = codec.encode(payload) + + /** The pre-#119 write plumbing, reconstructed: fresh BAOS + writer + encoder per call. */ + @Benchmark def naive_freshPlumbing: Array[Byte] = + val o = new ByteArrayOutputStream() + val writer = new GenericDatumWriter[Any](schema) + val enc = EncoderFactory.get().binaryEncoder(o, null) + writer.write(tree, enc) + enc.flush() + o.toByteArray + + @Benchmark def handwritten_stream: Array[Byte] = handwrittenToByteArray() + + private def handwrittenToByteArray(): Array[Byte] = + out.reset() + encoder = EncoderFactory.get().binaryEncoder(out, encoder) + val e = encoder + e.writeString(payload.id) + e.writeString(payload.tenant) + e.writeString(payload.source) + e.writeLong(payload.ts) + e.writeLong(payload.seq) + e.writeDouble(payload.amount) + e.writeBoolean(payload.flag) + e.writeInt(payload.kind) + payload.metrics match + case Some(m) => + e.writeIndex(metricsSomeIdx) + m.a1 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(0)); e.writeDouble(v) + case None => e.writeIndex(metricsNullIdx(0)) + m.a2 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(1)); e.writeDouble(v) + case None => e.writeIndex(metricsNullIdx(1)) + m.b1 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(2)); e.writeLong(v) + case None => e.writeIndex(metricsNullIdx(2)) + m.s1 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(3)); e.writeString(v) + case None => e.writeIndex(metricsNullIdx(3)) + m.i1 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(4)); e.writeInt(v) + case None => e.writeIndex(metricsNullIdx(4)) + m.b2 match + case Some(v) => e.writeIndex(1 - metricsNullIdx(5)); e.writeBoolean(v) + case None => e.writeIndex(metricsNullIdx(5)) + case None => e.writeIndex(nullArm(metricsUnion)) + payload.event match + case Event.Ev0(v) => e.writeIndex(0); e.writeLong(v) + case Event.Ev1(v) => e.writeIndex(1); e.writeLong(v) + case Event.Ev2(v) => e.writeIndex(2); e.writeLong(v) + case Event.Ev3(v) => e.writeIndex(3); e.writeLong(v) + case Event.Ev4(v) => e.writeIndex(4); e.writeLong(v) + case Event.Ev5(v) => e.writeIndex(5); e.writeLong(v) + e.flush() + out.toByteArray + end handwrittenToByteArray + +end AvroEncodeRouteBench diff --git a/docs/research/2026-09-29-issue-119-avro-encode-attribution.md b/docs/research/2026-09-29-issue-119-avro-encode-attribution.md new file mode 100644 index 00000000..7766feae --- /dev/null +++ b/docs/research/2026-09-29-issue-119-avro-encode-attribution.md @@ -0,0 +1,113 @@ +--- +date: 2026-09-29 +topic: issue-119-avro-encode-attribution +status: fix landed (per-thread write plumbing in `writeDatum`, route bench, safety pins) +scope: avro binary write path; plus a claim check on `generics` `lens[S]` focus ordering +--- + +# Issue #119 — decomposing the avro encode gap + +Trigger: [issue #119](https://github.com/Constructive-Programming/eo/issues/119) +reported that a kindlings/eo produce path (`AvroCodec.derived` → `AvroEncoder.derived` → +`GenericDatumWriter`) measures ~1.4–1.7× slower and 9–18% more allocated than a hand-written +direct-`BinaryEncoder` writer, and hypothesised the intermediate `GenericData.Record` tree. +The same report arrived in-thread bundled with a claim that `lens[S]` reorders its focus — +checked at the end; it does not. + +## The write path, decomposed + +A standalone attribution harness (published `cats-eo-avro` **0.17.0**, Scala 3.8.4, JDK 25, +`ThreadMXBean.getThreadAllocatedBytes`; fixture: 15-field record + 12-field all-`Option` +nested + 18-branch union, **245 B** output — the issue's shape at reduced scale; the +hand-written route was byte-identical to the eo route, so the comparison is honest): + +| Route | ns/op | B/op | +|-------------------------------------------------------------|------:|-----:| +| kindlings `encode(a)` → `Any` (tree only) | 532 | 664 | +| write datum with **fresh** BAOS + writer + encoder (pre-#119 `writeDatum`) | 1,599 | **2,928** | +| write datum with **reused** BAOS + encoder + cached writer | 1,311 | 392 | +| full eo route (`AvroCodec.encodeValue`) | 1,670 | 3,456 | +| hand-written fields straight to a reused `BinaryEncoder` | 761 | 397 | + +Two findings, and the reporter's hypothesis was only half right: + +1. **The time gap is tree + dispatch.** Tree build ≈ 1/3 of it; `GenericDatumWriter`'s + per-field walk (casts, position reads, `resolveUnion` per union field) ≈ the rest. That is + the *irreducible* cost of going through avro's datum model and lives in kindlings' + derivation, not in this repo. Reusing plumbing barely moves ns/op. +2. **The allocation gap was mostly ours.** `writeDatum` allocated a fresh + `ByteArrayOutputStream` (32-byte start — doubling and copying a dozen times toward a + multi-KB payload), a fresh `GenericDatumWriter`, and a fresh `BufferedBinaryEncoder` + (2 KB internal buffer) on **every call** — 2,928 B/op against a 392 B/op reused floor on a + 245 B record, scaling to ~2× the payload once the BAOS growth chain dominates. That is + the bulk of the reported 9–18% over a hand-written writer that already reuses its buffers. + +## The fix + +`AvroBinaryCursor.writeDatum` now delegates to a per-thread `DatumWriters` — the exact mirror +of what the READ side has done for a while (`DatumReaders` per-thread reader cache + +reusable `binaryDecoderCache`): + +- one `ByteArrayOutputStream` + one `BufferedBinaryEncoder` per writing thread, re-bound via + `EncoderFactory.binaryEncoder(out, reuse)` (avro's supported reconfigure path: position + reset, buffer kept); +- `GenericDatumWriter` cached per `Schema` in a `ThreadLocal` `HashMap` — the same shape and + the same thread-safety reasoning as the reader cache; +- `out.reset()` at the START of each encode and `toByteArray` at the end: every result is a + fresh detached array (retention-safe), and a mid-write throw leaves no observable residue. + +All five call sites (root `encodeRecord`/`encodeValue`, leaf-span splice, `.fields` overlay, +the circe + jsoniter bridge writes) funnel through this one choke point. + +Measured after the change (same harness): full eo route **3,456 → 904 B/op** (−74%); the +write side now costs exactly the returned array (~245 B + tree 664 B). Byte-identical output +against the old plumbing, asserted in-harness and covered by the module's exact-bytes specs. + +Pins (`AvroWriteCorrectnessSpec`): retention under 200 same-thread writes, clean recovery +after an aborted write, 8-thread concurrent encode equals single-writer golden bytes. + +## Re-measured permanently (`AvroEncodeRouteBench`) + +`sbt "benchmarks/Jmh/run -i 5 -wi 3 -f 3 -t 1 -prof gc .*AvroEncodeRouteBench.*"` — four +routes (tree-only, full `encodeValue`, the pre-fix fresh-plumbing shape kept as `naive_*` +baseline, hand-written stream) on a reduced version of the issue's record. Local quick profile +(one fork, `-prof gc`, gc.alloc.rate.norm — the repo's B/op-is-the-gate doctrine): + +| Route | ns/op | B/op | +|--------------------------|------:|-----:| +| `eo_encodeToAny` | 58 | 312 | +| `naive_freshPlumbing` | 744 | 2,624 | +| `eo_encodeValue` (fixed) | 758 | 568 | +| `handwritten_stream` | 174 | 240 | + +What's left between `eo_encodeValue` and `handwritten_stream` is the tree (≈ 312 B) and +`GenericDatumWriter` dispatch (time). Closing THAT gap is the encode-side twin of the +whole-record-builder work (#95): a derived `A => Encoder => Unit` streaming encoder — which +lives in kindlings-avro-derivation (its `AvroEncoder` typeclass is `A => Any`, so no +macro-side shim in eo can skip the datum). Worth an upstream ask; until then, +byte-carrying optics (`graftBytes`, offset-walk writes) remain the zero-tree escape hatch, +and the reported producer-vs-handwritten ALLOCATION delta is effectively gone on this side. + +## The bundled lens claim — not reproducible + +Claim: "`lens[Person](_.age, _.name)` should give focus `("age" -> Int, "name" -> String)` +but instead produces `("name" -> String, "age" -> Int)`." Checked against published +**cats-eo-generics 0.14.0** (the release current when the report was drafted) and **0.17.0**, +and against `HEAD`: + +- a `BijectionIso[Person, Person, NamedTuple[("age","name"), (Int,String)], …]` ascription + on `lens[Person](_.age, _.name)` **compiles** on both versions (all `Optic` parameters are + invariant, so a declaration-order focus could not typecheck); +- the compiler's own revealed type for that call: + `BijectionIso[Person, Person, (age : Int, name : String), …]` — selector order; +- a partial-cover pin, `SimpleLens[Employee, NamedTuple[("salary","id"), …], …]` for + `lens[Employee](_.salary, _.id)` (declaration order would be `("id","salary")`), also + compiles; named field access + laws run correctly at runtime; +- source-level: `buildMultiLens` / `buildMultiIso` have fed `namedTupleTypeOf(selectedNames, + …)` in selector order since the macro exists (v0.14.0 included), and `GenericsSpec` + exercises both the reversed and declaration-order selector spellings. + +The `("name" -> String, "age" -> Int)` display shape is a *dealiased* NamedTuple printing +(`Map[Values, Labels]` pairs); the current contract prints as `(age : Int, name : String)`. +If a future reporter can attach the exact version and the code that produced that display, +re-open — as filed, the claim does not reproduce.