Skip to content

avro: kindlings-derived AvroCodec encode is ~1.7× slower and allocates ~9–18% more than writing the same record straight to a BinaryEncoder #119

Description

Context

We replaced a vulcan Codec.toBinary on a hot produce path with two alternatives, both producing byte-identical Avro binary (verified by property tests against the vulcan bytes):

  1. Hand-written: fields written directly to an org.apache.avro.io.BinaryEncoder in schema order (writeIndex / writeString / writeLong …). Union payloads with ~18 branches still go through cached GenericDatumWriters.
  2. eo / kindlings: AvroEncoder.derived (via AvroCodec.derived) builds the Avro generic record, and a plain Avro GenericDatumWriter serializes it under the registered writer schema. Custom leaf encoders handle micros instants and map key order. A nested sub-record is spliced in with AvroPrism.graftBytes.

Both replace the same vulcan path. The record is a ~15-field top level with a nested ~55-field all-Option metrics record, an 18-branch union, and a ~20-field nested record that is either spliced (graft) or null.

Numbers (JMH, -prof gc, 1 fork, 5 warmup + 8 measurement iterations, µs/op and B/op)

Route vulcan hand-written eo / kindlings
with graft splice 6.85 µs, 22,272 B 1.05 µs, 6,904 B 1.81 µs, 7,528 B
without graft 3.26 µs, 12,328 B 1.01 µs, 3,856 B 1.40 µs, 4,544 B

The eo route is a big win over vulcan (3.8× / 2.3×), but it's about 1.4–1.7× slower than the hand-written writer and allocates 9–18% more. Caveat: the hand-written and eo numbers come from separate runs on the same machine; allocation should be reliable, timings indicative.

Hypothesis

The gap is the intermediate generic-record tree. AvroEncoder.derived materializes a GenericData.Record (plus boxed leaves / java.util collections) for every nested record, which GenericDatumWriter then walks with per-field schema dispatch. The hand-written path streams straight to the encoder.

Questions / possible directions

  • Could kindlings/eo derive a streaming encoder (A => BinaryEncoder => Unit) that writes fields in schema order without building the generic tree, keeping AvroCodec[A] as the typeclass? That's roughly the encode-side counterpart of what No efficient whole-record construction primitive (build-from-scratch beats AvroRecordPrism) #95 addressed for construction.
  • If not, is there a recommended pattern for writing a derived record directly to bytes (e.g. an AvroCodec method returning Array[Byte] that skips GenericData.Record)?
  • Is the extra allocation expected from the derived encoders (boxing Option / numeric leaves), or is it avoidable?

Versions: cats-eo-avro 0.14.0, kindlings-avro-derivation 0.3.2, hearth 0.4.2, avro 1.12.1, Scala 3.9.0, JDK 25 (Corretto).

Happy to share a reduced JMH reproducer if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions