Skip to content

Parquet Writer Optimization - #230

Merged
mchav merged 2 commits into
DataHaskell:mainfrom
sharmrj:parquet-writer-optimize-hot-loop
Sep 2, 2026
Merged

Parquet Writer Optimization#230
mchav merged 2 commits into
DataHaskell:mainfrom
sharmrj:parquet-writer-optimize-hot-loop

Conversation

@sharmrj

@sharmrj sharmrj commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

The hot loop in the Parquet Writer repeatedly calls writeRow which does a bunch of IORef bookkeeping. A lot of that book keeping can happen in between subbatches, which are already exposed via the write options. We refactor Encoder.hs, Writer.hs and use some new helper functions in RandomAccess.hs to make the hot loop faster by replacing writeRow with writeRows.

Results

Old

benchmarking write 10 GiB dataframe
time                 23.27 s    (23.04 s .. 23.45 s)
                     1.000 R²   (1.000 R² .. 1.000 R²)
mean                 23.29 s    (23.25 s .. 23.32 s)
std dev              41.82 ms   (19.60 ms .. 57.89 ms)
variance introduced by outliers: 19% (moderately inflated)

New

benchmarking write 10 GiB dataframe
time                 18.73 s    (18.70 s .. 18.77 s)
                     1.000 R²   (1.000 R² .. 1.000 R²)
mean                 18.84 s    (18.80 s .. 18.91 s)
std dev              71.08 ms   (10.62 ms .. 92.98 ms)
variance introduced by outliers: 19% (moderately inflated)

@mchav
mchav merged commit 000f111 into DataHaskell:main Sep 2, 2026
2 of 15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants