A byte-aligned binary format for Go structs, built for wires that are latency-bound: many small messages, a hot path, a known type on both ends.
data, err := colbin.Marshal(&charge)
err = colbin.Unmarshal(data, &back)Against protocol buffers on the same six-field record, protobuf driven through its generated code:
| one flat record | protobuf | colbin | |
|---|---|---|---|
| encode, reusing a buffer | 123 ns | 39 ns | 3.2× |
| encode, onto a fresh buffer | 140 ns | 76 ns | 1.8× |
| decode | 116 ns | 61 ns | 1.9× |
| bytes | 32 | 27 |
| one order, three nested lines | protobuf | colbin | |
|---|---|---|---|
| encode | 271 ns | 81 ns | 3.3× |
| decode | 485 ns | 314 ns | 1.5× |
| bytes | 49 | 46 |
colbin through Codec[T], protobuf through its generated code. Generating the
colbin codec instead takes the flat record to 27 ns encode and 55 ns
decode — 4.5× and 2.1×.
go test ./bench -bench . on an i7-1355U, Go 1.27, best of eight in one run.
packed5 off. Run a benchmark alone and it lands 5–10% faster than it does in the
sweep; both sides are measured the same way, so the ratios hold either way.
Nothing is packed across a byte boundary, and no size is a varint. A header carries the common size and, when it does not fit, names the width of the one that follows — so no read is ever a loop whose trip count is data, and a string field is a sub-slice of the message rather than a copy.
A field holding its zero value is not written at all, which is where most of the saving comes from.
There is one format. There used to be three modes and a byte at the front to tell them apart; that byte is now the root value's own descriptor, and it says the class and the key width.
A struct with no tags at all works: each field takes fnv8 of its name, linear
probed past anything already used. That id lands anywhere in 0..255, so an
untagged type uses eight-bit keys — which costs a byte per present field and
buys Skip over an unknown one.
type Charge struct {
CompanyID int32 // id = fnv8("CompanyID")
Name string
}Numbering the fields is how a type asks for the four-bit key, and it is the
number that goes on the wire, so it is what a reader in another language needs.
colbin.FieldIDs(v) prints whatever a type resolved to.
type Charge struct {
CompanyID int32 `cb:"0"`
UserID int32 `cb:"1"`
RouteID uint16 `cb:"2"`
Name string `cb:"3"`
}
data, err := colbin.Marshal(&charge)Codec[T] resolves the plan once and allocates nothing per record.
var chargeCodec = colbin.MustCodec[Charge]()
buf := make([]byte, 0, 64)
for _, charge := range charges {
buf = chargeCodec.Append(buf[:0], &charge)
send(buf)
}codec.Generate emits the straight-line calls, which is about three times
faster than the reflective walk:
| ten-field record | encode | decode |
|---|---|---|
hand-written against wire |
5.1 ns | 14.8 ns |
| generated | 8.8 ns | 21.1 ns |
Codec[T] handle |
22.3 ns | 25.3 ns |
Marshal / Unmarshal |
50.9 ns | 43.3 ns |
Taken in one run; absolute figures move ±20% between runs on this machine, so compare rows against each other rather than against a number taken elsewhere.
source, _ := codec.GenerateString("billing", Charge{})
os.WriteFile("charge_colbin.go", []byte(source), 0o644)A field id is four bits or eight, chosen per key run rather than per message.
| cost | buys | |
|---|---|---|
| 4-bit | — | the fast path: 5.3 ns encode, 14.8 ns decode |
| 8-bit | a byte per present field | 256 ids, Skip over an unknown field, packed5 |
Same six-field sensor reading, same run, protobuf through its generated code:
| encode | decode | bytes | |
|---|---|---|---|
| protobuf | 123 ns | 116 ns | 32 |
| colbin + tags (4-bit keys) | 40.6 ns | 62.3 ns | 27 |
| colbin untagged (8-bit keys) | 44.2 ns | 70.5 ns | 33 |
| colbin + packed5 (8-bit keys) | 52.1 ns | 69.4 ns | 33 |
Untagged costs 9% on encode and 13% on decode, and six bytes — one per present field. All of that is the key width, not the hashing, which happens once when the plan is built.
packed5 on this record is pure loss: its only string is "C". On a record that
plays to it — a product with a SKU, a name and three category strings — it is
still close to a wash:
| string-heavy product | encode | decode | bytes |
|---|---|---|---|
| protobuf | 170 ns | 389 ns | 81 |
| colbin + tags | 43.1 ns | 230 ns | 81 |
| colbin + packed5 | 282 ns | 427 ns | 80 |
One byte, for six times the encode cost. The packing saves 7 bytes and the wide key it forces costs 6. Turn it on only when strings dominate the record and size matters more than speed.
They are optimised separately, because what is scarce differs.
K4 has four descriptor bits and a reader that already knows the type. An
unsigned field therefore spends no sign bit: all sixteen codes carry
information, 0..7 being the value itself with no payload at all and 8..15 a
magnitude of one to eight bytes. A bool, a small count or a flag is a single
byte, key included, and the widths are exact — a seven-byte magnitude costs
seven where the signed form still rounds it to eight.
K8 has already spent a byte on the key, so its descriptor starts empty. It
gets a second form: a varint that carries three value bits in the descriptor and
seven per byte after it. The writer emits it only when it is shorter than the
sign-and-magnitude form, so nothing on the wire ever got bigger — a random
int64 still takes the byte count, because seven bits per byte loses to eight
once a value is wide.
| average bytes per field | sign+magnitude | with varint |
|---|---|---|
| small ids 0..1000 | 3.60 | 3.34 |
| deltas −1000..1000 | 3.67 | 3.42 |
| random int32 / int64 | 5.99 / 9.99 | 5.99 / 9.99 |
A signed field zigzags into the varint and an unsigned one does not, which is
safe for the same reason the K4 split is: the schema picks the reader. An
unknown field stays skippable either way — the varint is self-delimiting and
lives under a class, which is what Skip walks.
A type goes wide when it has an id above fifteen, or SetPacked5 on and a
string to spend it on. A nested struct, slice of structs, map or table does
not force it: a narrow descriptor carries the byte length those need, with the
class coming from the schema. Nothing else changes, and a
narrow-keyed struct can hold a wide-keyed one or the reverse — the width is in
the descriptor that opens each run.
A narrow message cannot skip an unknown field. Four descriptor bits have no room for a class, so a reader that does not know a key cannot size it. Adding a field is a coordinated deploy of both sides unless the type is on the wide path.
colbin.SetPacked5(true) turns on a string encoding worth about five bits per
character on upper-case alphanumerics. It is off by default and it is a
writer setting: the encoding is recorded in each string's own descriptor, so a
decoder reads either form without being told.
The encoder chooses per string, so turning it on can never make a message larger. It costs a pass over every string on both sides, and it puts the type on the wide key path.
| yes | bool, every sized int/uint, float32/64, string, []byte, slices of integers and of strings, nested structs, []struct, recursive types, map with string or integer keys, pointers to any scalar or string |
| not yet | interface{}, arrays, pointers to composites, maps of structs |
int and uint encode as their 64-bit forms, so a message written on one
platform reads on another.
A nil pointer is omitted and costs nothing. A non-nil pointer to a zero value
— new(int32), a *string to "" — writes an explicit zero, two bytes, because
otherwise it would be indistinguishable from nil. That is the one value in the
format written solely to say it is there.
type Patch struct {
Name *string `cb:"0"` // nil: leave it alone. &"": clear it.
Limit *int32 `cb:"1"`
}Pointers to structs, slices and maps are refused: those carry a length already, and what a nil one should mean is not settled.
Past a threshold, a []Struct field is transposed into a table: one key per
column rather than one per field per row, with each column through the blocked
column codec.
| six-field row | |
|---|---|
| list of structs, 7 rows | 16.9 B/row |
| table, 1000 rows | 11.0 B/row |
The choice is made per field on the row count, and the two are different descriptor classes, so a reader dispatches on what it finds. A struct with a nested struct or a slice inside it cannot be a column and stays row-wise however long it gets.
The wire carries no type — that is where the speed comes from — so a browser, a
jq-style tool or any dynamically typed client cannot name a field or tell a
float from an integer. A schema section gives it those. It is the same plan
the encoder already resolves from the struct, written out as bytes: a key, a
name and a type code per field, with nested structs hoisted into an indexed
table so a recursive type describes itself in finite space.
Send it once per connection, then send ordinary messages:
schema, _ := colbin.SchemaFor[Sale]()
send(schema.Bytes()) // once
for _, sale := range sales {
data, _ := colbin.Marshal(&sale) // unchanged, and unchanged in size
send(data)
}and on the other side:
schema, _ := colbin.ParseSchema(section)
text, _ := colbin.ToJSON(schema, message) // {"ID":1,"UserID":42,...}
value, _ := colbin.DecodeAny(schema, message)colbin.MarshalSelfDescribing(&sale) puts the section in front of the body
instead, for a document that has to stand alone. Its root byte is 0xD4 or
0xDC — the schema bit, 0x04 — and Unmarshal steps over the section, so a
self-describing message still decodes into the Go type.
It is the wrong default for a stream. Measured on the corpus:
| table | schema | B/message | schema/msg |
|---|---|---|---|
| users | 61 B | 61.9 B | 1.0x |
| sales (with detail) | 173 B | 89.5 B | 1.9x |
| metrics | 28 B | 10.8 B | 2.6x |
The JSON is what encoding/json would have written for the same record, down to
the escaping and the spelling of numbers — with two exceptions the wire forces:
an empty slice or map is indistinguishable from a nil one and comes out null,
and a NaN or an infinity is refused rather than quietly written as null.
DecodeAny keeps them.
Going straight to text is also the faster direction, because the intermediate
map[string]any is where all the allocation is:
| 100 corpus users | ns/op | B/op |
|---|---|---|
colbin.AppendJSON |
29 500 | 6 512 |
encoding/json on the structs |
40 000 | 14 327 |
colbin.DecodeAny |
41 000 | 54 712 |
Writing colbin from JSON is not in this: it needs type inference, and it is a separate job.
wire/ the format: field framing, all three key-run framings,
composites, tables, opt-in packed5
column/ the column codec: blocks of 128 residuals at a chosen bit width
codec/ the reflection façade and the source generator
packed5/ the opt-in string packing
corpus/ a reproducible, real-shaped dataset: users, products, sales
bench/ the comparison against protocol buffers
corpus.Generate(corpus.Seed, corpus.Small) builds the same seven tables every
time — users, products, categories, stores, sales with nested Detail []SaleLine, events and metrics. Money is integer cents throughout; a sale holds
no string and no float.
The line count per sale straddles the table threshold on purpose, so one dataset reaches both layouts:
| 300 sales, 1 514 lines | sales | lines | B/line |
|---|---|---|---|
| list of structs (<8 lines) | 246 | 820 | 21.9 |
| table, transposed (≥8) | 54 | 694 | 12.8 |
bench/corpus.pb.go is the protobuf twin, field for field. Cents are int64
rather than sint64 because every amount is non-negative and int64 is the
shorter of the two — protobuf gets its best form, not the matching one.
| table | rows | protobuf | colbin | |
|---|---|---|---|---|
| users | 100 | 6 275 | 6 187 | −1.4% |
| products | 200 | 12 996 | 12 876 | −0.9% |
| sales (nested detail) | 300 | 33 489 | 26 856 | −19.8% |
| metrics | 2 000 | 22 000 | 21 680 | −1.5% |
| total | 74 760 | 67 599 | −9.6% |
Per row, in one run:
| protobuf | colbin | ||
|---|---|---|---|
| user encode | 141 ns | 35 ns | 4.0× |
| user decode | 244 ns | 86 ns | 2.9× |
| sale encode | 616 ns | 363 ns | 1.7× |
| sale decode | 967 ns | 504 ns | 1.9× |
| metric encode | 61 ns | 15 ns | 4.0× |
Decoding a sale allocates 5.1 times against protobuf's 9.1.
Note the shape of the size result: on flat records the two formats are within 1.5% of each other — both omit zero fields and write a key per present field, so there is little to choose between them. The whole of colbin's size advantage is in the nested table, where a slice of integer-only structs is transposed into columns and protobuf has no equivalent.
go test ./corpus -run Report -v prints bytes per row for every table;
go test ./bench -run CorpusSizes -v prints the comparison above.
wire and column have no reflection and no type registry — they are driven by
a caller that already knows the Go type, which is what codec.Generate emits.
An integer column is a transform — raw, delta, frame-of-reference or constant —
then blocks of 128 residuals packed at an exact bit width chosen per block. 128
values at w bits is exactly 16w bytes, so a block is byte-aligned at both
ends and no state crosses a boundary.
| 256 × int64 | raw | encoded |
|---|---|---|
| monotonic ids | 2048 B | 171 B |
| timestamps | 2048 B | 235 B |
| all zeros | 2048 B | 3 B |
| random | 2048 B | 2051 B |
1.3 ns per element to decode, 3.5 to encode.
BYTE_ALIGNED_PLAN.md— the design, its measurements, and what is still openRATIONALE.md— the decisions, including the ones the measurements reversed and the optimisations that did not paywire/README.md,column/README.md— the layoutsrust/README.md— the Rust port
rust/ is the same format in Rust: the wire at all three key framings, the
column codec, packed5, and a #[derive(Colbin)] that emits the straight-line
encode and decode rather than a reflective walk.
#[derive(Colbin)]
struct Charge {
#[cb(0)] company_id: u32,
#[cb(1)] note: String,
}The two ports are pinned to each other rather than to a description.
rust/vectors/main.go writes a corpus with the Go codecs — the messages, the
field ids and the columns — and the Rust tests assert both directions against
it, so neither side can move without the other failing.
go run ./rust/vectors && go test ./rust/vectors
cargo test -p colbin --features deriveAlpha. The wire format is settled, the Go façade covers everything in the table above, and the Rust port covers the same ground. Interfaces have no form on the wire yet.
The schema section is Go-only so far: it changes nothing about the bytes an ordinary message carries, so the Rust port reads and writes those unaffected, but it cannot yet produce or consume a section of its own.