Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 33 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -60,6 +60,39 @@ true until the next version shipped.

No bash suite changes, so no ledger row moves and the census does not.
`cluster_tests` 418 -> 421, re-derived by collection.
- An encoding was chosen on pre-codec bytes but the chunk is stored post-codec,
so an encoding that shrank the bytes could enlarge the stored chunk (#1132).

`PgColumnarEncodeChunk` picks the smallest candidate against `bestLen`, which
starts at `rawLen`, and every comparison is on UNCOMPRESSED bytes. The block
codec runs afterwards, once, over the whole encoded region, defaulting to zstd
level 3. Bit-packing whitens a stream the codec was exploiting, so the two
disagree -- and only the codec's answer is what gets written.

FSST already decided this way, through `PgColumnarFsstHelpsCompressed`. The
writer now asks the same question for the rest: it compresses the encoded
region and the raw one and keeps whichever is smaller, per column chunk, which
is the granularity the codec actually runs at.

Measured on ClickBench `hits_0.parquet`, 1,000,000 rows and 105 columns,
imported with `pgcolumnar.import_parquet` on PG17:

stored total 81,869,112 -> 78,109,810 -4.59%
NONE vectors 101 -> 1,621
RLE / DICT / FOR 5372 / 3782 / 934 -> 4558 / 3390 / 621
FSST 310 -> 310 unchanged

FSST is unchanged because it already had this gate; the encoders that did not
are exactly the ones that moved. The worst single column, `ClientEventTime`,
was stored 49.7% smaller. Its shape is why: rare outliers stretch the range
frame-of-reference must size every value for, while the typical value's high
bytes stay constant for the codec to compress. Row counts, two column sums and
an md5 over `URL` are identical across the two loads.

`pgcolumnar.enable_post_codec_encoding_choice` (default `on`) restores the old
behaviour. It exists because the suites that test the ENCODERS need them to
actually run: a fixture chosen to exercise frame-of-reference packing is not
necessarily one where packing beats the codec.

- A covering projection was priced by clauses that merely mention its sort key,
rather than by clauses it can prune on (#1126, the remainder of #1107).
Expand Down
1 change: 1 addition & 0 deletions docs/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,7 @@ pgColumnar has two kinds of settings:
| `pgcolumnar.compression` | enum | `zstd` | Default block codec for new chunks. One of `none`, `pglz`, `lz4`, `zstd`. `lz4` and `zstd` are available only when the extension was built with those libraries. **This setting also participates in the lightweight encoding decisions that run before the codec**, so `none` is not the same cascade with compression removed; see [Compression and the encoding cascade](administration.md#compression-and-the-encoding-cascade). |
| `pgcolumnar.compression_level` | integer | `3` | Level for the `zstd` codec. Range 1 to 22. Higher levels compress more and write more slowly. |
| `pgcolumnar.fsst_min_gain_percent` | integer | `5` | Minimum size reduction, in percent, for FSST string encoding to be kept for a column chunk. Range 0 to 99. See below. |
| `pgcolumnar.enable_post_codec_encoding_choice` | boolean | `on` | Keep a lightweight encoding only when the encoded column chunk is smaller than the raw one AFTER the block codec has run. The encoders choose on uncompressed bytes, but what is stored is compressed, and bit-packing can whiten a stream the codec was exploiting; measured on ClickBench `hits`, 16 of 77 fixed-width columns were stored larger encoded than raw. Turning this off restores the pre-1132 behaviour of trusting the pre-codec choice. See [Compression and the encoding cascade](administration.md#compression-and-the-encoding-cascade). |
| `pgcolumnar.fsst_verdict_reuse` | integer | `16` | How many later row groups may reuse a column's FSST keep-or-drop verdict before it is decided again. Range 0 to INT_MAX. |
| `pgcolumnar.parallel_flush` | boolean | `off` | Opt-in. When on, a stripe flush of two or more columns fans the per-column encode and compress work out to background workers. The stored bytes match the serial path. It helps one large flush of many numeric columns by up to 14 percent. A wide text-heavy flush regresses, because it copies the buffered bytes through shared memory. Frequent small flushes regress too, so it is off by default. Enable it for a wide numeric bulk load in the session that runs it. |

Expand Down
1 change: 1 addition & 0 deletions src/columnar.h
Original file line number Diff line number Diff line change
Expand Up @@ -185,6 +185,7 @@ extern int pgcolumnar_fsst_verdict_reuse;
#define COLUMNAR_FSST_HELPS 1
#define COLUMNAR_FSST_HURTS 2
extern bool pgcolumnar_enable_qual_pushdown;
extern bool pgcolumnar_enable_post_codec_encoding_choice;
extern int pgcolumnar_qual_skipvec_min_payload_cols; /* #595 width gate */
extern bool pgcolumnar_enable_late_materialization;
extern bool pgcolumnar_enable_column_projection;
Expand Down
23 changes: 23 additions & 0 deletions src/columnar_tableam.c
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,7 @@ int pgcolumnar_compression_level = 3;
int pgcolumnar_fsst_min_gain_percent = 5;
int pgcolumnar_qual_skipvec_min_payload_cols = 20; /* #595 width gate */
bool pgcolumnar_enable_qual_pushdown = true;
bool pgcolumnar_enable_post_codec_encoding_choice = true;
bool pgcolumnar_enable_late_materialization = true;
bool pgcolumnar_enable_column_projection = true;
bool pgcolumnar_enable_bloom_filter = true;
Expand Down Expand Up @@ -3301,6 +3302,28 @@ _PG_init(void)
0,
NULL, NULL, NULL);

/*
* #1132. An encoding is chosen on pre-codec bytes but the chunk is stored
* post-codec, so the writer compares the encoded region against the raw one,
* both compressed, and keeps the smaller. Off restores the pre-#1132
* behaviour of trusting the pre-codec choice.
*
* It exists because the suites that test the ENCODERS need them to actually
* run: a fixture chosen to exercise frame-of-reference packing is not
* necessarily one where packing beats the codec, and those suites assert
* that the packer ran (encode_invariants.sh:139). Separating the two lets
* each be tested for what it is -- the encoders here, the selection policy
* in encode_post_codec.sh.
*/
DefineCustomBoolVariable("pgcolumnar.enable_post_codec_encoding_choice",
"Keep an encoding only when it is smaller after block compression.",
NULL,
&pgcolumnar_enable_post_codec_encoding_choice,
true,
PGC_USERSET,
0,
NULL, NULL, NULL);

/*
* Dev control for #393: off maps every page read 1:1 onto a ranged request
* so the request-count suite can measure both arms in one run. Nothing
Expand Down
93 changes: 93 additions & 0 deletions src/columnar_write_state.c
Original file line number Diff line number Diff line change
Expand Up @@ -1108,6 +1108,8 @@ flush_one_column(Form_pg_attribute att, List *chunkGroups,
uint64 rowIdx = 0;
StringInfo encoded = makeStringInfo();
StringInfo desc = makeStringInfo();
StringInfo rawRegion = makeStringInfo(); /* #1132: the unencoded alternative */
StringInfo rawDesc = makeStringInfo();
uint32 vectorCount = (uint32) list_length(chunkGroups);
char *fsstTable = NULL; /* chunk-shared FSST table (E3b), or NULL */
uint32 fsstTableLen = 0;
Expand Down Expand Up @@ -1157,6 +1159,7 @@ flush_one_column(Form_pg_attribute att, List *chunkGroups,

/* descriptor header (columnar_encdesc.h owns the wire layout) */
PgColumnarEncdescPutHeader(desc, vectorCount);
PgColumnarEncdescPutHeader(rawDesc, vectorCount);

/*
* E3b: build one FSST symbol table for the whole column chunk from a
Expand Down Expand Up @@ -1324,6 +1327,28 @@ flush_one_column(Form_pg_attribute att, List *chunkGroups,
PgColumnarEncdescPutEntry(desc, entryType, entryValueCount,
entryRawLen, encLen);

/*
* #1132: carry the unencoded alternative alongside. Encoding is chosen
* per vector on PRE-codec bytes, but the chunk is stored POST-codec, and
* bit-packing whitens a stream the codec was exploiting -- so a vector
* that shrank can still enlarge the stored chunk. Building both here
* lets the decision be made once, below, at the granularity the codec
* actually runs at.
*
* BUILT ONLY WHEN THE DECISION WILL BE TAKEN. This is a full copy of the
* column chunk's value stream, and what a flush holds is already a
* sensitivity here -- see the codec-buffer note below, measured on a
* 200,000-row load in #1075. With the choice off the copy is not made
* and the GUC costs nothing rather than costing memory silently.
*/
if (pgcolumnar_enable_post_codec_encoding_choice)
{
appendBinaryStringInfo(rawRegion, col->valueStream.data,
col->valueStream.len);
PgColumnarEncdescPutEntry(rawDesc, COLUMNAR_ENCODING_NONE,
entryValueCount, entryRawLen, entryRawLen);
}

/* per-vector zone map (native spec 7.1, D5) */
{
NativeZoneMapMetadata *z = palloc0(sizeof(NativeZoneMapMetadata));
Expand Down Expand Up @@ -1395,6 +1420,18 @@ flush_one_column(Form_pg_attribute att, List *chunkGroups,
if (fsstTableLen > 0)
appendBinaryStringInfo(desc, fsstTable, fsstTableLen);

/*
* The unencoded alternative needs the same trailing region, but never a
* table: its entries are all COLUMNAR_ENCODING_NONE, so nothing can
* reference one. Writing the length unconditionally keeps the exact-length
* check in columnar_reader.c:1111-1115 satisfied either way.
*/
{
uint32 noSharedTable = 0;

appendBinaryStringInfo(rawDesc, (char *) &noSharedTable, sizeof(uint32));
}

/* whole-chunk zone map (vector_index -1) */
{
NativeZoneMapMetadata *z = palloc0(sizeof(NativeZoneMapMetadata));
Expand Down Expand Up @@ -1487,6 +1524,62 @@ flush_one_column(Form_pg_attribute att, List *chunkGroups,
finalLen = compLen;
blockCodec = usedType;
}

/*
* #1132: THE ENCODING IS CHOSEN PRE-CODEC AND THE CHUNK IS STORED
* POST-CODEC, so ask the only question that decides the stored size --
* is the encoded region, once compressed, actually smaller than the raw
* one compressed? Measured on ClickBench hits_0.parquet, the answer was
* no on 16 of 77 fixed-width columns, costing 7.17% of their stored
* bytes, and ClientEventTime alone was stored 2.06x larger than raw.
*
* FSST already decides this way, through PgColumnarFsstHelpsCompressed.
* This is the same question asked for the rest.
*
* ONLY WHEN ENCODING CLAIMED A WIN. `rawRegion->len > encoded->len` is
* the cheap precondition: when the encoders all declined, the two
* regions are the same bytes and compressing twice would buy nothing.
* It also bounds the added cost to chunks where there is a decision to
* make.
*/
if (pgcolumnar_enable_post_codec_encoding_choice &&
rawRegion->len > encoded->len)
{
char *rawCodecBuf = NULL;
uint32 rawCompLen;
int rawUsedType;
int rawUsedLevel;
uint32 rawFinalLen;

PgColumnarCompressValueStream(rawRegion->data, rawRegion->len,
compressionType,
compressionLevel,
&rawCodecBuf, &rawCompLen,
&rawUsedType, &rawUsedLevel);
rawFinalLen = (rawUsedType != COLUMNAR_COMPRESSION_NONE)
? rawCompLen : rawRegion->len;

if (rawFinalLen < finalLen)
{
/*
* Storing it unencoded wins. The descriptor must describe the
* bytes actually written, so swap it for the all-NONE one built
* alongside; a descriptor that disagrees with its chunk is a
* decode error, not a size regression.
*/
if (codecBuf != NULL)
pfree(codecBuf);
codecBuf = rawCodecBuf;
finalData = (rawUsedType != COLUMNAR_COMPRESSION_NONE)
? rawCodecBuf : rawRegion->data;
finalLen = rawFinalLen;
blockCodec = (rawUsedType != COLUMNAR_COMPRESSION_NONE)
? rawUsedType : COLUMNAR_COMPRESSION_NONE;
desc = rawDesc;
}
else
pfree(rawCodecBuf);
}
}

if (finalLen > 0)
Expand Down
11 changes: 11 additions & 0 deletions test/check_ledger.tsv
Original file line number Diff line number Diff line change
Expand Up @@ -202,6 +202,17 @@ differential differential wide point 15;16;17;18;19 never -
differential differential wide range 15;16;17;18;19 never -
differential differential wide row proj 15;16;17;18;19 never -
differential differential wide row scan 15;16;17;18;19 never -
encode_post_codec encode_post_codec a chunk is not stored larger than it would be with no encoding at all 15;16;17;18;19 never -
encode_post_codec encode_post_codec and that column stays far below the no-encoding size, so the win is real 15;16;17;18;19 never -
encode_post_codec encode_post_codec premise: both fixtures hold every row 15;16;17;18;19 never -
encode_post_codec encode_post_codec premise: the tail fixture's range is set by outliers, not by its typical value 15;16;17;18;19 never -
encode_post_codec encode_post_codec the rep column holds no row the heap does not 15;16;17;18;19 never -
encode_post_codec encode_post_codec the rep column preserves its checksum 15;16;17;18;19 never -
encode_post_codec encode_post_codec the rep column reads back exactly what the heap holds 15;16;17;18;19 never -
encode_post_codec encode_post_codec the tail column holds no row the heap does not 15;16;17;18;19 never -
encode_post_codec encode_post_codec the tail column preserves its checksum 15;16;17;18;19 never -
encode_post_codec encode_post_codec the tail column reads back exactly what the heap holds 15;16;17;18;19 never -
encode_post_codec encode_post_codec while a column where encoding genuinely wins still encodes 15;16;17;18;19 never -
harness_selftest 030-assertions nothing leaked into the squatter 15;16;17;18;19 never -
harness_selftest 030-assertions pgc_port_free says the squatter's port is busy 15;16;17;18;19 never -
harness_selftest 030-assertions squatter survived untouched 15;16;17;18;19 never -
Expand Down
2 changes: 1 addition & 1 deletion test/check_ledger_budget.txt
Original file line number Diff line number Diff line change
Expand Up @@ -128,4 +128,4 @@ suites_not_covered 249
# one short. The census command above was never affected because it skips nothing,
# but a row count taken the other way disagrees with the tool's own `ledger: rows=`
# and reads as an off-by-one in the merge rather than in the command.
checks_never_observed_red 1396
checks_never_observed_red 1407
9 changes: 9 additions & 0 deletions test/encode_invariants.sh
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,15 @@ set -uo pipefail
. "$(dirname "${BASH_SOURCE[0]}")/lib.sh"
pgc_setup "${1:-/usr/local/pg17/bin/pg_config}"

# #1132: the writer now keeps an encoding only when it is smaller AFTER block
# compression, and this suite's subject is the ENCODERS rather than that choice.
# Its fixtures are chosen to exercise a particular encoder, which is not the same
# as being fixtures where that encoder beats zstd -- two of its controls assert
# the encoder actually ran, and those went red when the choice landed. Pinning the
# pre-#1132 behaviour here keeps this suite testing what it is named for; the
# selection policy has its own suite, encode_post_codec.sh.
psql_run "ALTER DATABASE $PGC_DB SET pgcolumnar.enable_post_codec_encoding_choice = off;" >/dev/null 2>&1

ROWS="${PGC_ENCINV_ROWS:-4096}"

# The C-level half. Bound here rather than shipped, like the other debug hooks.
Expand Down
Loading
Loading