Skip to content

CBMC: Refine bounds for input and output of base multiplication - #906

Draft
hanno-becker wants to merge 1 commit into
mainfrom
refine_bounds
Draft

CBMC: Refine bounds for input and output of base multiplication#906
hanno-becker wants to merge 1 commit into
mainfrom
refine_bounds

Conversation

@hanno-becker

@hanno-becker hanno-becker commented Mar 24, 2025

Copy link
Copy Markdown
Contributor

Previously, the base multiplication would assume that one of its inputs is bound by 4096 in absolute value, but make no assumptions about the other input ("b-input" henceforth) and its mulcache.

This commit refines the bounds slightly, as follows:

  • The b-input is assumed to be bound by MLK_NTT_BOUND in absolute value. This comes for free since all values for b are results of the NTT.
  • The b-cache-input is assumed to be bound by MLKEM_Q in absolute value.

With those additional bounds in place, it can be showed that the result of the base multiplication is below INT16_MAX/2 in absolute value. Accordingly, this can be added as a precondition for the inverse NTT.

For the native AVX2 backend, the new output bound for the mulcache forces an explicit zeroization of the mulcache. This is not ideal since the cache is in fact entirely unused, but the performance penalty should be marginal (if the compiler can't eliminate the zeroization in the first place).

@hanno-becker hanno-becker added enhancement New feature or request CBMC labels Mar 24, 2025
@hanno-becker
hanno-becker force-pushed the refine_bounds branch 4 times, most recently from 1326e42 to f9b25ee Compare March 24, 2025 11:12
@hanno-becker
hanno-becker marked this pull request as ready for review March 24, 2025 12:56
@hanno-becker
hanno-becker requested a review from a team as a code owner March 24, 2025 12:56
@hanno-becker hanno-becker added the benchmark this PR should be benchmarked in CI label Mar 24, 2025

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Mac Mini (M1, 2020) benchmarks

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 12257 cycles 12254 cycles 1.00
ML-KEM-512 encaps 14848 cycles 14846 cycles 1.00
ML-KEM-512 decaps 19725 cycles 19721 cycles 1.00
ML-KEM-768 keypair 21048 cycles 21051 cycles 1.00
ML-KEM-768 encaps 23677 cycles 23679 cycles 1.00
ML-KEM-768 decaps 30515 cycles 30515 cycles 1
ML-KEM-1024 keypair 29977 cycles 29978 cycles 1.00
ML-KEM-1024 encaps 34479 cycles 34478 cycles 1.00
ML-KEM-1024 decaps 43430 cycles 43442 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 11882 cycles 11814 cycles 1.01
ML-KEM-512 encaps 13289 cycles 13091 cycles 1.02
ML-KEM-512 decaps 17239 cycles 17198 cycles 1.00
ML-KEM-768 keypair 19720 cycles 19378 cycles 1.02
ML-KEM-768 encaps 21078 cycles 20699 cycles 1.02
ML-KEM-768 decaps 26624 cycles 26393 cycles 1.01
ML-KEM-1024 keypair 28312 cycles 28230 cycles 1.00
ML-KEM-1024 encaps 30253 cycles 30160 cycles 1.00
ML-KEM-1024 decaps 37612 cycles 37619 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 4th gen (c7i) (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 27656 cycles 27528 cycles 1.00
ML-KEM-512 encaps 34464 cycles 34189 cycles 1.01
ML-KEM-512 decaps 44045 cycles 43887 cycles 1.00
ML-KEM-768 keypair 44260 cycles 44314 cycles 1.00
ML-KEM-768 encaps 54929 cycles 54874 cycles 1.00
ML-KEM-768 decaps 68271 cycles 68308 cycles 1.00
ML-KEM-1024 keypair 67587 cycles 67436 cycles 1.00
ML-KEM-1024 encaps 78819 cycles 78816 cycles 1.00
ML-KEM-1024 decaps 96111 cycles 96150 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A55 (Snapdragon 888) benchmarks

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 59336 cycles 59330 cycles 1.00
ML-KEM-512 encaps 66928 cycles 66920 cycles 1.00
ML-KEM-512 decaps 86001 cycles 86014 cycles 1.00
ML-KEM-768 keypair 101044 cycles 101087 cycles 1.00
ML-KEM-768 encaps 112086 cycles 112317 cycles 1.00
ML-KEM-768 decaps 139218 cycles 139358 cycles 1.00
ML-KEM-1024 keypair 153559 cycles 153471 cycles 1.00
ML-KEM-1024 encaps 172905 cycles 170285 cycles 1.02
ML-KEM-1024 decaps 210032 cycles 207958 cycles 1.01

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 12844 cycles 12818 cycles 1.00
ML-KEM-512 encaps 14281 cycles 14213 cycles 1.00
ML-KEM-512 decaps 19056 cycles 18981 cycles 1.00
ML-KEM-768 keypair 21662 cycles 21611 cycles 1.00
ML-KEM-768 encaps 22851 cycles 22673 cycles 1.01
ML-KEM-768 decaps 29760 cycles 29642 cycles 1.00
ML-KEM-1024 keypair 30561 cycles 30460 cycles 1.00
ML-KEM-1024 encaps 32758 cycles 32619 cycles 1.00
ML-KEM-1024 decaps 42208 cycles 42024 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 14506 cycles 14430 cycles 1.01
ML-KEM-512 encaps 16067 cycles 15953 cycles 1.01
ML-KEM-512 decaps 21519 cycles 21352 cycles 1.01
ML-KEM-768 keypair 23895 cycles 23653 cycles 1.01
ML-KEM-768 encaps 25261 cycles 25052 cycles 1.01
ML-KEM-768 decaps 33298 cycles 32878 cycles 1.01
ML-KEM-1024 keypair 33747 cycles 33529 cycles 1.01
ML-KEM-1024 encaps 36060 cycles 35883 cycles 1.00
ML-KEM-1024 decaps 46499 cycles 46302 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 18012 cycles 17951 cycles 1.00
ML-KEM-512 encaps 20181 cycles 20119 cycles 1.00
ML-KEM-512 decaps 26910 cycles 26773 cycles 1.01
ML-KEM-768 keypair 30869 cycles 30083 cycles 1.03
ML-KEM-768 encaps 32097 cycles 31885 cycles 1.01
ML-KEM-768 decaps 44369 cycles 41521 cycles 1.07
ML-KEM-1024 keypair 42410 cycles 42120 cycles 1.01
ML-KEM-1024 encaps 45721 cycles 45813 cycles 1.00
ML-KEM-1024 decaps 58408 cycles 58278 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Performance Alert ⚠️

Possible performance regression was detected for benchmark 'Intel Xeon 3rd gen (c6i)'.
Benchmark result of this commit is worse than the previous benchmark result exceeding threshold 1.03.

Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-768 decaps 44369 cycles 41521 cycles 1.07

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 4th gen (c7a) (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 36736 cycles 36732 cycles 1.00
ML-KEM-512 encaps 42809 cycles 42788 cycles 1.00
ML-KEM-512 decaps 55544 cycles 55534 cycles 1.00
ML-KEM-768 keypair 58381 cycles 58385 cycles 1.00
ML-KEM-768 encaps 67032 cycles 67035 cycles 1.00
ML-KEM-768 decaps 83945 cycles 84011 cycles 1.00
ML-KEM-1024 keypair 88541 cycles 88488 cycles 1.00
ML-KEM-1024 encaps 99207 cycles 99002 cycles 1.00
ML-KEM-1024 decaps 120514 cycles 120405 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AMD EPYC 3rd gen (c6a) (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 40213 cycles 39898 cycles 1.01
ML-KEM-512 encaps 48345 cycles 48282 cycles 1.00
ML-KEM-512 decaps 61864 cycles 61774 cycles 1.00
ML-KEM-768 keypair 62689 cycles 62604 cycles 1.00
ML-KEM-768 encaps 75020 cycles 74835 cycles 1.00
ML-KEM-768 decaps 92705 cycles 92457 cycles 1.00
ML-KEM-1024 keypair 95554 cycles 95266 cycles 1.00
ML-KEM-1024 encaps 110497 cycles 110126 cycles 1.00
ML-KEM-1024 decaps 133215 cycles 132700 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 17664 cycles 17687 cycles 1.00
ML-KEM-512 encaps 20548 cycles 20545 cycles 1.00
ML-KEM-512 decaps 27003 cycles 27028 cycles 1.00
ML-KEM-768 keypair 29873 cycles 29839 cycles 1.00
ML-KEM-768 encaps 32701 cycles 32723 cycles 1.00
ML-KEM-768 decaps 41923 cycles 41863 cycles 1.00
ML-KEM-1024 keypair 43715 cycles 43694 cycles 1.00
ML-KEM-1024 encaps 48640 cycles 48696 cycles 1.00
ML-KEM-1024 decaps 61390 cycles 61378 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intel Xeon 3rd gen (c6i) (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 46949 cycles 46967 cycles 1.00
ML-KEM-512 encaps 55709 cycles 55697 cycles 1.00
ML-KEM-512 decaps 71513 cycles 71394 cycles 1.00
ML-KEM-768 keypair 74347 cycles 74347 cycles 1
ML-KEM-768 encaps 85920 cycles 85936 cycles 1.00
ML-KEM-768 decaps 106934 cycles 106925 cycles 1.00
ML-KEM-1024 keypair 111324 cycles 111271 cycles 1.00
ML-KEM-1024 encaps 126039 cycles 126002 cycles 1.00
ML-KEM-1024 decaps 152116 cycles 152097 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton4 (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 35215 cycles 35180 cycles 1.00
ML-KEM-512 encaps 40026 cycles 39967 cycles 1.00
ML-KEM-512 decaps 50570 cycles 50600 cycles 1.00
ML-KEM-768 keypair 56693 cycles 56654 cycles 1.00
ML-KEM-768 encaps 64089 cycles 63984 cycles 1.00
ML-KEM-768 decaps 78446 cycles 78490 cycles 1.00
ML-KEM-1024 keypair 87373 cycles 87513 cycles 1.00
ML-KEM-1024 encaps 96554 cycles 96618 cycles 1.00
ML-KEM-1024 decaps 114862 cycles 114683 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 18864 cycles 18877 cycles 1.00
ML-KEM-512 encaps 22344 cycles 22358 cycles 1.00
ML-KEM-512 decaps 29577 cycles 29566 cycles 1.00
ML-KEM-768 keypair 32236 cycles 32250 cycles 1.00
ML-KEM-768 encaps 35628 cycles 35598 cycles 1.00
ML-KEM-768 decaps 45941 cycles 45965 cycles 1.00
ML-KEM-1024 keypair 46491 cycles 46487 cycles 1.00
ML-KEM-1024 encaps 52069 cycles 52080 cycles 1.00
ML-KEM-1024 decaps 65916 cycles 65962 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 29410 cycles 29444 cycles 1.00
ML-KEM-512 encaps 34613 cycles 34680 cycles 1.00
ML-KEM-512 decaps 45213 cycles 45279 cycles 1.00
ML-KEM-768 keypair 50051 cycles 50139 cycles 1.00
ML-KEM-768 encaps 55328 cycles 55224 cycles 1.00
ML-KEM-768 decaps 70158 cycles 70096 cycles 1.00
ML-KEM-1024 keypair 73082 cycles 73098 cycles 1.00
ML-KEM-1024 encaps 81472 cycles 81514 cycles 1.00
ML-KEM-1024 decaps 101379 cycles 101471 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton3 (no-opt)

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 38966 cycles 39685 cycles 0.98
ML-KEM-512 encaps 44939 cycles 44864 cycles 1.00
ML-KEM-512 decaps 56734 cycles 56842 cycles 1.00
ML-KEM-768 keypair 64234 cycles 64182 cycles 1.00
ML-KEM-768 encaps 72031 cycles 72669 cycles 0.99
ML-KEM-768 decaps 88070 cycles 88056 cycles 1.00
ML-KEM-1024 keypair 96228 cycles 96095 cycles 1.00
ML-KEM-1024 encaps 106307 cycles 106211 cycles 1.00
ML-KEM-1024 decaps 127120 cycles 127057 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

SpacemiT K1 8 (Banana Pi F3) benchmarks

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 225601 cycles 225593 cycles 1.00
ML-KEM-512 encaps 271839 cycles 271821 cycles 1.00
ML-KEM-512 decaps 346406 cycles 346326 cycles 1.00
ML-KEM-768 keypair 374348 cycles 374216 cycles 1.00
ML-KEM-768 encaps 433877 cycles 433736 cycles 1.00
ML-KEM-768 decaps 532770 cycles 532650 cycles 1.00
ML-KEM-1024 keypair 554398 cycles 554463 cycles 1.00
ML-KEM-1024 encaps 632515 cycles 632521 cycles 1.00
ML-KEM-1024 decaps 754467 cycles 754358 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A72 (Raspberry Pi 4) benchmarks

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 52821 cycles 53196 cycles 0.99
ML-KEM-512 encaps 60537 cycles 61239 cycles 0.99
ML-KEM-512 decaps 77877 cycles 77578 cycles 1.00
ML-KEM-768 keypair 89670 cycles 89929 cycles 1.00
ML-KEM-768 encaps 97185 cycles 98860 cycles 0.98
ML-KEM-768 decaps 121486 cycles 123436 cycles 0.98
ML-KEM-1024 keypair 135647 cycles 134458 cycles 1.01
ML-KEM-1024 encaps 148998 cycles 147089 cycles 1.01
ML-KEM-1024 decaps 180865 cycles 180649 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton2 (no-opt)

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 59326 cycles 59431 cycles 1.00
ML-KEM-512 encaps 68033 cycles 68011 cycles 1.00
ML-KEM-512 decaps 86704 cycles 86649 cycles 1.00
ML-KEM-768 keypair 98815 cycles 98924 cycles 1.00
ML-KEM-768 encaps 110974 cycles 110402 cycles 1.01
ML-KEM-768 decaps 134574 cycles 134811 cycles 1.00
ML-KEM-1024 keypair 148680 cycles 148822 cycles 1.00
ML-KEM-1024 encaps 163724 cycles 163875 cycles 1.00
ML-KEM-1024 decaps 195449 cycles 195498 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Arm Cortex-A76 (Raspberry Pi 5) benchmarks

Details
Benchmark suite Current: ee52c4e Previous: 8dba13b Ratio
ML-KEM-512 keypair 29409 cycles 29440 cycles 1.00
ML-KEM-512 encaps 34618 cycles 34675 cycles 1.00
ML-KEM-512 decaps 45210 cycles 45264 cycles 1.00
ML-KEM-768 keypair 50053 cycles 50138 cycles 1.00
ML-KEM-768 encaps 55322 cycles 55220 cycles 1.00
ML-KEM-768 decaps 70154 cycles 70075 cycles 1.00
ML-KEM-1024 keypair 73076 cycles 73079 cycles 1.00
ML-KEM-1024 encaps 81445 cycles 81517 cycles 1.00
ML-KEM-1024 decaps 101376 cycles 101473 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@hanno-becker hanno-becker removed the benchmark this PR should be benchmarked in CI label Mar 24, 2025
@hanno-becker

Copy link
Copy Markdown
Contributor Author

@dkostic Would you (and Claude) have time to do this, and at the same time resolve #1421?

@rod-chapman

Copy link
Copy Markdown
Contributor

I hope that CBMC 6.9.0 will unblock progress on this one. I will rebase the branch first and see if the proofs complete OK.

@hanno-becker

Copy link
Copy Markdown
Contributor Author

Thanks @rod-chapman! But this also needs significant work on the x86_64 backend.

@rod-chapman

Copy link
Copy Markdown
Contributor

Oh yes... the x86 back-end has to brought into line with the AArch64 and C back-ends to produce tighter bounds on the output of basemul, right?

I have re-based and re-running proofs now on the C stuff.

@oqs-bot

oqs-bot commented May 4, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-512)

⚠️ Attention Required

Proof Status Current Previous Change
**TOTAL** ⚠️ 2225s 1614s +37.9%
mlk_fqmul ⚠️ 29s 17s +71%
mlk_indcpa_dec ⚠️ 36s 14s +157%
mlk_indcpa_enc ⚠️ 467s 255s +83%
mlk_indcpa_keypair_derand ⚠️ 502s 213s +136%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 132s 55s +140%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 50s 8s +525%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** ⚠️ 2225s 1614s +37.9%
mlk_indcpa_keypair_derand ⚠️ 502s 213s +136%
mlk_indcpa_enc ⚠️ 467s 255s +83%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 132s 55s +140%
mlk_rej_uniform_c 113s 135s -16%
rej_uniform_native_x86_64 52s 57s -9%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 50s 8s +525%
rej_uniform_native_aarch64 49s 43s +14%
poly_ntt_native 45s 45s +0%
mlk_poly_reduce_native 42s 35s +20%
mlk_poly_rej_uniform 41s 150s -73%
mlk_ntt_layer 40s 32s +25%
mlk_indcpa_dec ⚠️ 36s 14s +157%
mlk_keccak_squeezeblocks_x4 30s 26s +15%
mlk_fqmul ⚠️ 29s 17s +71%
keccakf1600x4_permute_native_x4 17s 16s +6%
mlk_poly_decompress_d4_native 17s 14s +21%
mlk_poly_decompress_d10_native 15s 15s +0%
mlk_poly_frommsg 14s 10s +40%
mlk_poly_frombytes_native 11s 8s +38%
mlk_polyvec_add 11s 11s +0%
mlk_invntt_layer 10s 4s +150%
mlk_keccak_squeezeblocks 10s 9s +11%
mlk_keccak_squeeze_once 9s 8s +12%
mlk_matvec_mul 9s 3s +200%
mlk_ntt_butterfly_block 9s 9s +0%
mlk_poly_ntt 9s 5s +80%
kem_dec 8s 5s +60%
mlk_keccak_absorb_once_x4 7s 5s +40%
mlk_poly_rej_uniform_x4 7s 8s -12%
mlk_poly_sub 7s 6s +17%
mlk_polymat_permute_bitrev_to_custom 7s 2s +250%
poly_frombytes_native_x86_64 7s 5s +40%
mlk_poly_cbd_eta2 6s 3s +100%
poly_decompress_d10_native_x86_64 6s 9s -33%
kem_check_pk 5s 4s +25%
kem_check_sk 5s 2s +150%
mlk_poly_add 5s 4s +25%
mlk_poly_compress_d5_c 5s 3s +67%
mlk_poly_decompress_d5_c 5s 1s +400%
mlk_poly_getnoise_eta1_4x_native 5s 2s +150%
mlk_poly_mulcache_compute_c 5s 4s +25%
mlk_poly_tobytes_native 5s 2s +150%
mlk_polyvec_frombytes 5s 2s +150%
mlk_polyvec_invntt_tomont 5s 3s +67%
poly_decompress_d4_native_x86_64 5s 4s +25%
mlk_ct_cmask_neg_i16 4s 3s +33%
mlk_gen_matrix 4s 2s +100%
mlk_keccak_absorb_once 4s 3s +33%
mlk_keccakf1600_permute 4s 1s +300%
mlk_keccakf1600_permute_c 4s 5s -20%
mlk_keccakf1600_xor_bytes (big endian) 4s 2s +100%
mlk_keccakf1600x4_extract_bytes_c 4s 3s +33%
mlk_poly_frombytes_c 4s 4s +0%
mlk_poly_getnoise_eta2 4s 3s +33%
mlk_poly_invntt_tomont_c 4s 3s +33%
mlk_poly_reduce 4s 1s +300%
mlk_poly_tomsg 4s 3s +33%
mlk_scalar_compress_d10 4s 2s +100%
mlk_shake128_absorb_once 4s 3s +33%
mlk_shake128_squeezeblocks 4s 3s +33%
mlk_value_barrier_u32 4s 3s +33%
poly_decompress_d5_native_x86_64 4s 2s +100%
poly_invntt_tomont_native 4s 2s +100%
poly_reduce_native_aarch64 4s 1s +300%
poly_tobytes_native_x86_64 4s 1s +300%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 4s 3s +33%
intt_native_aarch64 3s 2s +50%
keccak_f1600_x1_native_aarch64_v84a 3s 2s +50%
keccak_f1600_x4_native_aarch64_v84a 3s 3s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
kem_enc 3s 2s +50%
kem_enc_derand 3s 3s +0%
kem_keypair 3s 4s -25%
mlk_check_pct 3s 5s -40%
mlk_ct_cmask_nonzero_u16 3s 2s +50%
mlk_ct_cmov_zero 3s 3s +0%
mlk_keccakf1600_extract_bytes 3s 2s +50%
mlk_keccakf1600_xor_bytes 3s 1s +200%
mlk_poly_compress_d10_c 3s 2s +50%
mlk_poly_compress_d4 3s 3s +0%
mlk_poly_compress_d4_native 3s 2s +50%
mlk_poly_compress_dv 3s 3s +0%
mlk_poly_decompress_d4_c 3s 7s -57%
mlk_poly_frombytes 3s 1s +200%
mlk_poly_mulcache_compute 3s 2s +50%
mlk_poly_reduce_c 3s 3s +0%
mlk_poly_tobytes_c 3s 3s +0%
mlk_poly_tomont_c 3s 2s +50%
mlk_polyvec_permute_bitrev_to_custom_native 3s 2s +50%
mlk_scalar_compress_d11 3s 2s +50%
mlk_scalar_compress_d4 3s 1s +200%
mlk_scalar_compress_d5 3s 2s +50%
mlk_scalar_decompress_d4 3s 3s +0%
mlk_scalar_decompress_d5 3s 3s +0%
mlk_shake128x4_absorb_once 3s 2s +50%
mlk_shake128x4_squeezeblocks 3s 1s +200%
mlk_shake256 3s 2s +50%
mlk_shake256x4 3s 5s -40%
mlk_value_barrier_i32 3s 2s +50%
nttunpack_native_x86_64 3s 2s +50%
poly_compress_d10_native_x86_64 3s 3s +0%
poly_compress_d4_native_x86_64 3s 1s +200%
poly_getnoise_eta1122_4x_native 3s 1s +200%
poly_tobytes_native_aarch64 3s 2s +50%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 3s 3s +0%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 3s 2s +50%
intt_native_x86_64 2s 2s +0%
keccak_f1600_x1_native_aarch64 2s 2s +0%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccakf1600_permute_native 2s 1s +100%
keccakf1600x4_extract_bytes_native 2s 1s +100%
keccakf1600x4_xor_bytes_native 2s 1s +100%
kem_keypair_derand 2s 3s -33%
mlk_ct_cmask_nonzero_u8 2s 1s +100%
mlk_ct_get_optblocker_i32 2s 2s +0%
mlk_ct_get_optblocker_u32 2s 3s -33%
mlk_ct_get_optblocker_u8 2s 2s +0%
mlk_ct_memcmp 2s 3s -33%
mlk_ct_sel_int16 2s 1s +100%
mlk_enc_getnoise_eta1_eta2 2s 2s +0%
mlk_gen_matrix_serial 2s 3s -33%
mlk_keccakf1600x4_extract_bytes 2s 2s +0%
mlk_keccakf1600x4_xor_bytes 2s 2s +0%
mlk_keccakf1600x4_xor_bytes_c 2s 2s +0%
mlk_keypair_getnoise_eta1 2s 1s +100%
mlk_montgomery_reduce 2s 2s +0%
mlk_poly_compress_d10 2s 4s -50%
mlk_poly_compress_d11_c 2s 2s +0%
mlk_poly_compress_du 2s 3s -33%
mlk_poly_decompress_d11_c 2s 2s +0%
mlk_poly_decompress_d11_native 2s 3s -33%
mlk_poly_decompress_d4 2s 3s -33%
mlk_poly_decompress_d5 2s 1s +100%
mlk_poly_decompress_du 2s 2s +0%
mlk_poly_decompress_dv 2s 2s +0%
mlk_poly_getnoise_eta1122_4x 2s 3s -33%
mlk_poly_getnoise_eta1_4x 2s 5s -60%
mlk_poly_ntt_c 2s 2s +0%
mlk_poly_tobytes 2s 3s -33%
mlk_poly_tomont 2s 1s +100%
mlk_poly_tomont_native 2s 2s +0%
mlk_polyvec_basemul_acc_montgomery_cached 2s 1s +100%
mlk_polyvec_compress_du 2s 2s +0%
mlk_polyvec_decompress_du 2s 2s +0%
mlk_polyvec_ntt 2s 1s +100%
mlk_polyvec_permute_bitrev_to_custom 2s 2s +0%
mlk_polyvec_reduce 2s 2s +0%
mlk_rej_uniform 2s 5s -60%
mlk_scalar_compress_d1 2s 2s +0%
mlk_scalar_decompress_d11 2s 2s +0%
mlk_sha3_512 2s 2s +0%
mlk_value_barrier_u8 2s 4s -50%
poly_compress_d11_native_x86_64 2s 2s +0%
poly_compress_d5_native_x86_64 2s 2s +0%
poly_decompress_d11_native_x86_64 2s 1s +100%
poly_mulcache_compute_native_aarch64 2s 3s -33%
poly_mulcache_compute_native_x86_64 2s 1s +100%
poly_tomont_native_aarch64 2s 2s +0%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 2s 3s -33%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 2s 2s +0%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 2s 3s -33%
rej_uniform_native 2s 1s +100%
sys_check_capability 2s 1s +100%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 1s 3s -67%
mlk_barrett_reduce 1s 2s -50%
mlk_ct_sel_uint8 1s 2s -50%
mlk_keccakf1600_extract_bytes (big endian) 1s 5s -80%
mlk_keccakf1600x4_permute 1s 4s -75%
mlk_poly_cbd_eta1 1s 2s -50%
mlk_poly_compress_d10_native 1s 1s +0%
mlk_poly_compress_d11 1s 3s -67%
mlk_poly_compress_d11_native 1s 3s -67%
mlk_poly_compress_d4_c 1s 3s -67%
mlk_poly_compress_d5 1s 2s -50%
mlk_poly_compress_d5_native 1s 3s -67%
mlk_poly_decompress_d10 1s 2s -50%
mlk_poly_decompress_d10_c 1s 4s -75%
mlk_poly_decompress_d11 1s 2s -50%
mlk_poly_decompress_d5_native 1s 2s -50%
mlk_poly_invntt_tomont 1s 2s -50%
mlk_poly_mulcache_compute_native 1s 2s -50%
mlk_polyvec_mulcache_compute 1s 5s -80%
mlk_polyvec_tobytes 1s 3s -67%
mlk_polyvec_tomont 1s 1s +0%
mlk_scalar_decompress_d10 1s 2s -50%
mlk_scalar_signed_to_unsigned_q 1s 1s +0%
mlk_sha3_256 1s 2s -50%
ntt_native_aarch64 1s 2s -50%
ntt_native_x86_64 1s 2s -50%
poly_reduce_native_x86_64 1s 2s -50%
poly_tomont_native_x86_64 1s 2s -50%

@oqs-bot

oqs-bot commented May 4, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-768)

⚠️ Attention Required

Proof Status Current Previous Change
**TOTAL** ⚠️ 2367s 1645s +43.9%
kem_dec ⚠️ 27s 5s +440%
mlk_fqmul ⚠️ 24s 16s +50%
mlk_indcpa_dec ⚠️ 30s 10s +200%
mlk_indcpa_enc ⚠️ 430s 252s +71%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 566s 46s +1130%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 138s 14s +886%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** ⚠️ 2367s 1645s +43.9%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 566s 46s +1130%
mlk_indcpa_enc ⚠️ 430s 252s +71%
mlk_indcpa_keypair_derand 273s 318s -14%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 138s 14s +886%
mlk_rej_uniform_c 86s 114s -25%
rej_uniform_native_x86_64 45s 55s -18%
rej_uniform_native_aarch64 44s 45s -2%
poly_ntt_native 39s 39s +0%
mlk_poly_rej_uniform 35s 129s -73%
mlk_poly_reduce_native 33s 34s -3%
mlk_ntt_layer 31s 29s +7%
mlk_indcpa_dec ⚠️ 30s 10s +200%
kem_dec ⚠️ 27s 5s +440%
mlk_keccak_squeezeblocks_x4 27s 24s +12%
mlk_fqmul ⚠️ 24s 16s +50%
keccakf1600x4_permute_native_x4 16s 21s -24%
mlk_poly_decompress_d4_native 15s 12s +25%
mlk_polyvec_add 13s 12s +8%
mlk_poly_decompress_d10_native 12s 15s -20%
mlk_poly_frommsg 10s 9s +11%
mlk_keccak_squeeze_once 8s 7s +14%
mlk_keccak_squeezeblocks 8s 8s +0%
mlk_ntt_butterfly_block 8s 7s +14%
mlk_poly_frombytes_native 8s 8s +0%
mlk_poly_rej_uniform_x4 8s 5s +60%
mlk_keccak_absorb_once_x4 7s 5s +40%
mlk_poly_ntt 7s 5s +40%
mlk_poly_sub 7s 10s -30%
mlk_poly_add 6s 7s -14%
poly_decompress_d11_native_x86_64 6s 1s +500%
kem_enc_derand 5s 3s +67%
mlk_invntt_layer 5s 4s +25%
mlk_keccakf1600_permute_c 5s 6s -17%
mlk_poly_frombytes_c 5s 4s +25%
mlk_poly_getnoise_eta1_4x 5s 2s +150%
mlk_poly_mulcache_compute 5s 2s +150%
mlk_poly_tomont_native 5s 3s +67%
mlk_polyvec_tobytes 5s 1s +400%
mlk_shake256x4 5s 3s +67%
poly_mulcache_compute_native_x86_64 5s 3s +67%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 5s 1s +400%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 4s 3s +33%
kem_check_pk 4s 2s +100%
kem_enc 4s 2s +100%
mlk_gen_matrix 4s 2s +100%
mlk_gen_matrix_serial 4s 2s +100%
mlk_poly_compress_d10_c 4s 2s +100%
mlk_poly_decompress_d11 4s 2s +100%
mlk_poly_invntt_tomont 4s 4s +0%
mlk_poly_ntt_c 4s 1s +300%
mlk_polyvec_compress_du 4s 1s +300%
mlk_scalar_compress_d11 4s 1s +300%
mlk_scalar_decompress_d4 4s 2s +100%
poly_decompress_d10_native_x86_64 4s 5s -20%
poly_decompress_d4_native_x86_64 4s 6s -33%
poly_frombytes_native_x86_64 4s 5s -20%
poly_getnoise_eta1122_4x_native 4s 1s +300%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 4s 2s +100%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 3s 1s +200%
kem_keypair 3s 2s +50%
mlk_check_pct 3s 3s +0%
mlk_ct_cmask_neg_i16 3s 3s +0%
mlk_ct_get_optblocker_u32 3s 1s +200%
mlk_enc_getnoise_eta1_eta2 3s 3s +0%
mlk_keccak_absorb_once 3s 3s +0%
mlk_keccakf1600x4_extract_bytes_c 3s 2s +50%
mlk_keccakf1600x4_permute 3s 4s -25%
mlk_keypair_getnoise_eta1 3s 2s +50%
mlk_matvec_mul 3s 4s -25%
mlk_poly_cbd_eta2 3s 2s +50%
mlk_poly_compress_d10 3s 2s +50%
mlk_poly_compress_d11_c 3s 1s +200%
mlk_poly_compress_d11_native 3s 4s -25%
mlk_poly_compress_d5_c 3s 1s +200%
mlk_poly_compress_dv 3s 1s +200%
mlk_poly_decompress_d4 3s 3s +0%
mlk_poly_decompress_d4_c 3s 4s -25%
mlk_poly_decompress_d5_c 3s 3s +0%
mlk_poly_decompress_du 3s 1s +200%
mlk_poly_getnoise_eta1122_4x 3s 1s +200%
mlk_poly_mulcache_compute_native 3s 1s +200%
mlk_poly_tobytes 3s 3s +0%
mlk_poly_tomont_c 3s 3s +0%
mlk_poly_tomsg 3s 3s +0%
mlk_polyvec_frombytes 3s 3s +0%
mlk_polyvec_mulcache_compute 3s 2s +50%
mlk_polyvec_reduce 3s 1s +200%
mlk_polyvec_tomont 3s 1s +200%
mlk_scalar_signed_to_unsigned_q 3s 5s -40%
mlk_shake128_absorb_once 3s 1s +200%
mlk_shake256 3s 2s +50%
mlk_value_barrier_u32 3s 2s +50%
poly_compress_d11_native_x86_64 3s 3s +0%
poly_tobytes_native_aarch64 3s 1s +200%
poly_tomont_native_aarch64 3s 1s +200%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 3s 3s +0%
sys_check_capability 3s 3s +0%
intt_native_aarch64 2s 5s -60%
intt_native_x86_64 2s 3s -33%
keccak_f1600_x1_native_aarch64 2s 1s +100%
keccak_f1600_x4_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccakf1600x4_extract_bytes_native 2s 3s -33%
keccakf1600x4_xor_bytes_native 2s 2s +0%
kem_check_sk 2s 2s +0%
mlk_ct_cmask_nonzero_u16 2s 4s -50%
mlk_ct_cmask_nonzero_u8 2s 4s -50%
mlk_ct_get_optblocker_i32 2s 1s +100%
mlk_ct_get_optblocker_u8 2s 2s +0%
mlk_ct_sel_int16 2s 3s -33%
mlk_keccakf1600_extract_bytes 2s 4s -50%
mlk_keccakf1600_extract_bytes (big endian) 2s 2s +0%
mlk_keccakf1600_permute 2s 1s +100%
mlk_keccakf1600_xor_bytes 2s 3s -33%
mlk_keccakf1600_xor_bytes (big endian) 2s 1s +100%
mlk_keccakf1600x4_extract_bytes 2s 4s -50%
mlk_keccakf1600x4_xor_bytes_c 2s 1s +100%
mlk_montgomery_reduce 2s 1s +100%
mlk_poly_cbd_eta1 2s 1s +100%
mlk_poly_compress_d11 2s 1s +100%
mlk_poly_compress_d4 2s 3s -33%
mlk_poly_compress_d4_c 2s 1s +100%
mlk_poly_compress_d5 2s 2s +0%
mlk_poly_compress_d5_native 2s 2s +0%
mlk_poly_compress_du 2s 1s +100%
mlk_poly_decompress_d10 2s 2s +0%
mlk_poly_decompress_d10_c 2s 1s +100%
mlk_poly_decompress_d5 2s 2s +0%
mlk_poly_decompress_d5_native 2s 3s -33%
mlk_poly_decompress_dv 2s 2s +0%
mlk_poly_getnoise_eta1_4x_native 2s 3s -33%
mlk_poly_getnoise_eta2 2s 2s +0%
mlk_poly_invntt_tomont_c 2s 2s +0%
mlk_poly_mulcache_compute_c 2s 2s +0%
mlk_poly_reduce 2s 2s +0%
mlk_poly_reduce_c 2s 3s -33%
mlk_poly_tobytes_c 2s 1s +100%
mlk_poly_tomont 2s 1s +100%
mlk_polymat_permute_bitrev_to_custom 2s 3s -33%
mlk_polyvec_basemul_acc_montgomery_cached 2s 4s -50%
mlk_polyvec_decompress_du 2s 4s -50%
mlk_polyvec_ntt 2s 4s -50%
mlk_polyvec_permute_bitrev_to_custom 2s 1s +100%
mlk_rej_uniform 2s 1s +100%
mlk_scalar_compress_d10 2s 2s +0%
mlk_scalar_compress_d4 2s 2s +0%
mlk_scalar_decompress_d11 2s 3s -33%
mlk_scalar_decompress_d5 2s 2s +0%
mlk_sha3_256 2s 2s +0%
mlk_shake128x4_absorb_once 2s 3s -33%
ntt_native_x86_64 2s 3s -33%
nttunpack_native_x86_64 2s 3s -33%
poly_compress_d10_native_x86_64 2s 5s -60%
poly_compress_d4_native_x86_64 2s 3s -33%
poly_compress_d5_native_x86_64 2s 2s +0%
poly_decompress_d5_native_x86_64 2s 4s -50%
poly_invntt_tomont_native 2s 2s +0%
poly_mulcache_compute_native_aarch64 2s 6s -67%
poly_reduce_native_aarch64 2s 2s +0%
poly_reduce_native_x86_64 2s 2s +0%
poly_tobytes_native_x86_64 2s 3s -33%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 2s 2s +0%
rej_uniform_native 2s 2s +0%
keccak_f1600_x1_native_aarch64_v84a 1s 3s -67%
keccakf1600_permute_native 1s 3s -67%
kem_keypair_derand 1s 2s -50%
mlk_barrett_reduce 1s 2s -50%
mlk_ct_cmov_zero 1s 3s -67%
mlk_ct_memcmp 1s 4s -75%
mlk_ct_sel_uint8 1s 1s +0%
mlk_keccakf1600x4_xor_bytes 1s 2s -50%
mlk_poly_compress_d10_native 1s 2s -50%
mlk_poly_compress_d4_native 1s 3s -67%
mlk_poly_decompress_d11_c 1s 3s -67%
mlk_poly_decompress_d11_native 1s 1s +0%
mlk_poly_frombytes 1s 2s -50%
mlk_poly_tobytes_native 1s 1s +0%
mlk_polyvec_invntt_tomont 1s 1s +0%
mlk_polyvec_permute_bitrev_to_custom_native 1s 2s -50%
mlk_scalar_compress_d1 1s 2s -50%
mlk_scalar_compress_d5 1s 1s +0%
mlk_scalar_decompress_d10 1s 2s -50%
mlk_sha3_512 1s 3s -67%
mlk_shake128_squeezeblocks 1s 3s -67%
mlk_shake128x4_squeezeblocks 1s 1s +0%
mlk_value_barrier_i32 1s 2s -50%
mlk_value_barrier_u8 1s 2s -50%
ntt_native_aarch64 1s 3s -67%
poly_tomont_native_x86_64 1s 1s +0%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 1s 4s -75%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 1s 2s -50%

@oqs-bot

oqs-bot commented May 4, 2026

Copy link
Copy Markdown
Contributor

CBMC Results (ML-KEM-1024)

⚠️ Attention Required

Proof Status Current Previous Change
**TOTAL** ⚠️ 2509s 1885s +33.1%
kem_dec ⚠️ 21s 5s +320%
mlk_indcpa_dec ⚠️ 35s 11s +218%
mlk_indcpa_enc ⚠️ 617s 402s +53%
mlk_indcpa_keypair_derand ⚠️ 364s 176s +107%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 203s 79s +157%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 373s 31s +1103%
Full Results (191 proofs)
Proof Status Current Previous Change
**TOTAL** ⚠️ 2509s 1885s +33.1%
mlk_indcpa_enc ⚠️ 617s 402s +53%
polyvec_basemul_acc_montgomery_cached_native ⚠️ 373s 31s +1103%
mlk_indcpa_keypair_derand ⚠️ 364s 176s +107%
mlk_polyvec_basemul_acc_montgomery_cached_c ⚠️ 203s 79s +157%
mlk_rej_uniform_c 81s 144s -44%
rej_uniform_native_x86_64 43s 67s -36%
rej_uniform_native_aarch64 42s 53s -21%
mlk_poly_rej_uniform 36s 161s -78%
mlk_indcpa_dec ⚠️ 35s 11s +218%
poly_ntt_native 35s 43s -19%
mlk_poly_reduce_native 32s 40s -20%
mlk_fqmul 28s 20s +40%
mlk_ntt_layer 27s 38s -29%
mlk_keccak_squeezeblocks_x4 25s 26s -4%
kem_dec ⚠️ 21s 5s +320%
mlk_polyvec_add 18s 18s +0%
keccakf1600x4_permute_native_x4 16s 15s +7%
mlk_polyvec_ntt 15s 12s +25%
mlk_poly_decompress_d11_native 14s 16s -12%
mlk_poly_decompress_d5_native 12s 18s -33%
mlk_keccak_squeezeblocks 9s 9s +0%
mlk_poly_frombytes_native 9s 10s -10%
mlk_poly_frommsg 9s 10s -10%
mlk_keccak_absorb_once_x4 8s 5s +60%
mlk_poly_sub 8s 8s +0%
mlk_keccak_squeeze_once 7s 10s -30%
mlk_keccakf1600_permute_c 7s 6s +17%
mlk_ntt_butterfly_block 7s 9s -22%
mlk_invntt_layer 6s 6s +0%
mlk_poly_ntt 6s 8s -25%
mlk_poly_rej_uniform_x4 6s 7s -14%
mlk_polyvec_invntt_tomont 6s 1s +500%
kem_enc 5s 2s +150%
mlk_gen_matrix 5s 2s +150%
mlk_keccak_absorb_once 5s 4s +25%
mlk_matvec_mul 5s 5s +0%
mlk_poly_add 5s 5s +0%
mlk_poly_frombytes_c 5s 3s +67%
nttunpack_native_x86_64 5s 2s +150%
polyvec_basemul_acc_montgomery_cached_k2_native_x86_64 5s 2s +150%
mlk_poly_compress_d10_native 4s 3s +33%
mlk_poly_compress_d11_c 4s 4s +0%
mlk_poly_compress_du 4s 2s +100%
mlk_poly_decompress_d11_c 4s 2s +100%
mlk_poly_decompress_d5_c 4s 2s +100%
mlk_poly_mulcache_compute 4s 2s +100%
mlk_poly_mulcache_compute_c 4s 3s +33%
mlk_poly_tobytes_native 4s 1s +300%
mlk_poly_tomont 4s 2s +100%
mlk_polyvec_basemul_acc_montgomery_cached 4s 2s +100%
mlk_shake256x4 4s 4s +0%
poly_compress_d10_native_x86_64 4s 2s +100%
poly_decompress_d4_native_x86_64 4s 2s +100%
polyvec_basemul_acc_montgomery_cached_k4_native_aarch64 4s 2s +100%
intt_native_aarch64 3s 3s +0%
keccak_f1600_x4_native_aarch64_v84a 3s 3s +0%
keccak_f1600_x4_native_aarch64_v8a_scalar_hybrid 3s 3s +0%
keccakf1600x4_extract_bytes_native 3s 2s +50%
kem_check_pk 3s 4s -25%
kem_keypair 3s 2s +50%
kem_keypair_derand 3s 1s +200%
mlk_check_pct 3s 1s +200%
mlk_ct_get_optblocker_i32 3s 3s +0%
mlk_ct_get_optblocker_u32 3s 3s +0%
mlk_ct_sel_int16 3s 4s -25%
mlk_gen_matrix_serial 3s 4s -25%
mlk_keccakf1600_extract_bytes 3s 2s +50%
mlk_keccakf1600_extract_bytes (big endian) 3s 2s +50%
mlk_poly_cbd_eta1 3s 3s +0%
mlk_poly_compress_d11_native 3s 3s +0%
mlk_poly_compress_d4 3s 3s +0%
mlk_poly_decompress_d10_c 3s 2s +50%
mlk_poly_decompress_d10_native 3s 2s +50%
mlk_poly_decompress_d4 3s 3s +0%
mlk_poly_decompress_d5 3s 3s +0%
mlk_poly_getnoise_eta1122_4x 3s 3s +0%
mlk_poly_getnoise_eta1_4x 3s 2s +50%
mlk_poly_tobytes_c 3s 4s -25%
mlk_poly_tomont_native 3s 4s -25%
mlk_poly_tomsg 3s 3s +0%
mlk_polyvec_permute_bitrev_to_custom_native 3s 6s -50%
mlk_polyvec_tomont 3s 2s +50%
mlk_scalar_compress_d5 3s 4s -25%
mlk_scalar_decompress_d11 3s 2s +50%
mlk_scalar_decompress_d5 3s 2s +50%
mlk_sha3_512 3s 1s +200%
poly_compress_d11_native_x86_64 3s 3s +0%
poly_compress_d4_native_x86_64 3s 2s +50%
poly_decompress_d11_native_x86_64 3s 5s -40%
poly_decompress_d5_native_x86_64 3s 5s -40%
poly_frombytes_native_x86_64 3s 5s -40%
poly_getnoise_eta1122_4x_native 3s 3s +0%
poly_invntt_tomont_native 3s 3s +0%
poly_mulcache_compute_native_aarch64 3s 3s +0%
poly_mulcache_compute_native_x86_64 3s 5s -40%
poly_reduce_native_aarch64 3s 2s +50%
polyvec_basemul_acc_montgomery_cached_k3_native_x86_64 3s 2s +50%
polyvec_basemul_acc_montgomery_cached_k4_native_x86_64 3s 2s +50%
intt_native_x86_64 2s 2s +0%
keccak_f1600_x1_native_aarch64_v84a 2s 2s +0%
keccak_f1600_x4_native_aarch64_v8a_v84a_scalar_hybrid 2s 3s -33%
keccak_f1600_x4_native_avx2 2s 2s +0%
keccakf1600_permute_native 2s 2s +0%
kem_check_sk 2s 3s -33%
kem_enc_derand 2s 3s -33%
mlk_ct_cmask_nonzero_u16 2s 2s +0%
mlk_ct_cmask_nonzero_u8 2s 1s +100%
mlk_ct_cmov_zero 2s 3s -33%
mlk_ct_get_optblocker_u8 2s 2s +0%
mlk_ct_memcmp 2s 3s -33%
mlk_ct_sel_uint8 2s 2s +0%
mlk_enc_getnoise_eta1_eta2 2s 4s -50%
mlk_keccakf1600_permute 2s 2s +0%
mlk_keccakf1600_xor_bytes (big endian) 2s 2s +0%
mlk_keccakf1600x4_extract_bytes_c 2s 3s -33%
mlk_keccakf1600x4_xor_bytes 2s 2s +0%
mlk_keypair_getnoise_eta1 2s 5s -60%
mlk_poly_compress_d10 2s 2s +0%
mlk_poly_compress_d10_c 2s 3s -33%
mlk_poly_compress_d11 2s 3s -33%
mlk_poly_compress_d4_c 2s 4s -50%
mlk_poly_compress_d5_c 2s 4s -50%
mlk_poly_compress_dv 2s 1s +100%
mlk_poly_decompress_d4_native 2s 3s -33%
mlk_poly_decompress_du 2s 1s +100%
mlk_poly_frombytes 2s 2s +0%
mlk_poly_getnoise_eta1_4x_native 2s 3s -33%
mlk_poly_invntt_tomont 2s 3s -33%
mlk_poly_invntt_tomont_c 2s 2s +0%
mlk_poly_mulcache_compute_native 2s 3s -33%
mlk_poly_ntt_c 2s 1s +100%
mlk_poly_tobytes 2s 3s -33%
mlk_poly_tomont_c 2s 3s -33%
mlk_polymat_permute_bitrev_to_custom 2s 5s -60%
mlk_polyvec_compress_du 2s 3s -33%
mlk_polyvec_decompress_du 2s 2s +0%
mlk_polyvec_frombytes 2s 5s -60%
mlk_polyvec_mulcache_compute 2s 4s -50%
mlk_polyvec_permute_bitrev_to_custom 2s 4s -50%
mlk_polyvec_reduce 2s 2s +0%
mlk_scalar_compress_d1 2s 3s -33%
mlk_scalar_compress_d10 2s 2s +0%
mlk_scalar_compress_d11 2s 4s -50%
mlk_scalar_decompress_d4 2s 3s -33%
mlk_sha3_256 2s 2s +0%
mlk_shake128_absorb_once 2s 2s +0%
mlk_shake128_squeezeblocks 2s 1s +100%
mlk_shake128x4_absorb_once 2s 2s +0%
mlk_value_barrier_i32 2s 3s -33%
mlk_value_barrier_u32 2s 3s -33%
mlk_value_barrier_u8 2s 1s +100%
ntt_native_aarch64 2s 3s -33%
poly_compress_d5_native_x86_64 2s 4s -50%
poly_reduce_native_x86_64 2s 3s -33%
poly_tobytes_native_aarch64 2s 3s -33%
poly_tobytes_native_x86_64 2s 3s -33%
poly_tomont_native_aarch64 2s 3s -33%
poly_tomont_native_x86_64 2s 2s +0%
polyvec_basemul_acc_montgomery_cached_k2_native_aarch64 2s 1s +100%
rej_uniform_native 2s 1s +100%
sys_check_capability 2s 2s +0%
keccak_f1600_x1_native_aarch64 1s 1s +0%
keccakf1600x4_xor_bytes_native 1s 2s -50%
mlk_barrett_reduce 1s 2s -50%
mlk_ct_cmask_neg_i16 1s 3s -67%
mlk_keccakf1600_xor_bytes 1s 1s +0%
mlk_keccakf1600x4_extract_bytes 1s 2s -50%
mlk_keccakf1600x4_permute 1s 2s -50%
mlk_keccakf1600x4_xor_bytes_c 1s 5s -80%
mlk_montgomery_reduce 1s 2s -50%
mlk_poly_cbd_eta2 1s 3s -67%
mlk_poly_compress_d4_native 1s 2s -50%
mlk_poly_compress_d5 1s 5s -80%
mlk_poly_compress_d5_native 1s 4s -75%
mlk_poly_decompress_d10 1s 2s -50%
mlk_poly_decompress_d11 1s 3s -67%
mlk_poly_decompress_d4_c 1s 4s -75%
mlk_poly_decompress_dv 1s 2s -50%
mlk_poly_getnoise_eta2 1s 1s +0%
mlk_poly_reduce 1s 2s -50%
mlk_poly_reduce_c 1s 1s +0%
mlk_polyvec_tobytes 1s 1s +0%
mlk_rej_uniform 1s 2s -50%
mlk_scalar_compress_d4 1s 1s +0%
mlk_scalar_decompress_d10 1s 2s -50%
mlk_scalar_signed_to_unsigned_q 1s 3s -67%
mlk_shake128x4_squeezeblocks 1s 2s -50%
mlk_shake256 1s 1s +0%
ntt_native_x86_64 1s 2s -50%
poly_decompress_d10_native_x86_64 1s 4s -75%
polyvec_basemul_acc_montgomery_cached_k3_native_aarch64 1s 2s -50%

@rod-chapman

Copy link
Copy Markdown
Contributor

Update. I am having trouble getting one proof to work - it's polvec_basemul_acc_montgomery_cached_native().
The problem is the need to carry the stronger pre- and post-conditions across the call to the native back-end in a way that they "survive" so that, if the first call to the native backend responds with "FALLBACK", they are still true for the second call to the C backend. The problem is that the native backend changes representation of the types (e.g. from polyvec * to int16_t b[MLKEM_K * MLKEM_N which makes the reasoning sufficiently complex to defeat Z3.

@rod-chapman
rod-chapman force-pushed the refine_bounds branch 2 times, most recently from 9e4db99 to 7f2d2bc Compare May 8, 2026 14:43
@rod-chapman

Copy link
Copy Markdown
Contributor

C Proofs are OK now. I found a Z3 tactic that was effective. @dkostic The next move is to prove that the bounds on the outputs of the AArch64 native implementation of basemul are as expected, and as given by the specification of those 3 functions in native/aarch64/src/arith_native_aarch64.h on the refine_bounds branch. The CBMC spec of that for K=2 is here:

ensures(array_abs_bound(r, 0, MLKEM_N, INT16_MAX/2))

The critical contract is ensures(array_abs_bound(r, 0, MLKEM_N, INT16_MAX/2))

We need to know that's true for the assembly language versions (3 times for K=2,3,4), so that means amending the HOL-Light proof to explicitly state that post-condition. For K=2, that's here or thereabouts:

==> (ival(read(memory :> bytes16(word_add dst (word (4 * i)))) s) ==

Does that make sense?

@rod-chapman
rod-chapman force-pushed the refine_bounds branch 2 times, most recently from 085d1f0 to 999f564 Compare May 18, 2026 07:45
@dkostic
dkostic force-pushed the refine_bounds branch 4 times, most recently from 457c655 to bc59d2d Compare June 12, 2026 19:04
@dkostic

dkostic commented Jun 22, 2026

Copy link
Copy Markdown
Contributor

Before and after performance measured on m8i EC2 instance with Intel(R) Xeon(R) 6975P-C:

  ┌──────────────┬──────────────────┬────────┬────────┬──────────────────┬────────┬────────┬───────────────────┬─────────┬─────────┐
  │              │ mlkem512 keypair │ encaps │ decaps │ mlkem768 keypair │ encaps │ decaps │ mlkem1024 keypair │ encaps  │ decaps  │
  ├──────────────┼──────────────────┼────────┼────────┼──────────────────┼────────┼────────┼───────────────────┼─────────┼─────────┤
  │ no Barrett   │ 37 359           │ 43 892 │ 55 961 │ 60 371           │ 68 916 │ 84 209 │ 90 615            │ 104 190 │ 122 639 │
  ├──────────────┼──────────────────┼────────┼────────┼──────────────────┼────────┼────────┼───────────────────┼─────────┼─────────┤
  │ with Barrett │ 37 345           │ 43 784 │ 55 847 │ 60 867           │ 69 213 │ 84 345 │ 90 648            │ 104 279 │ 122 820 │
  ├──────────────┼──────────────────┼────────┼────────┼──────────────────┼────────┼────────┼───────────────────┼─────────┼─────────┤
  │ Δ cycles     │ −14              │ −108   │ −114   │ +496             │ +297   │ +136   │ +33               │ +89     │ +181    │
  ├──────────────┼──────────────────┼────────┼────────┼──────────────────┼────────┼────────┼───────────────────┼─────────┼─────────┤
  │ Δ %          │ −0.04%           │ −0.25% │ −0.20% │ +0.82%           │ +0.43% │ +0.16% │ +0.04%            │ +0.09%  │ +0.15%  │
  └──────────────┴──────────────────┴────────┴────────┴──────────────────┴────────┴────────┴───────────────────┴─────────┴─────────┘

Tighten the base-multiplication output bounds across the CBMC and
HOL Light proofs by constraining the second operand and its
multiplication cache.

AArch64:
- HOL Light: prove output bounds for poly_basemul_acc_montgomery_cached
  k2/k3/k4 under |b| < 8*MLKEM_Q and |b_cache| < MLKEM_Q, giving
  abs(r) <= 8323 / 11652 / 14981 respectively (previously the looser
  9857 / 13953 / 18049 under unconstrained b/b_cache).
- Add matching CBMC input/output bound contracts to the basemul
  declarations.

x86_64:
- Append a Barrett-reduction sweep to the AVX2 basemul kernels so every
  output coefficient is bounded by abs(r) < 2*MLKEM_Q regardless of k
  (previously grew to ~8*Q at k=4).
- HOL Light: update the basemul proofs for the Barrett reduction
  (refresh bytecode, fold barred_x86, hoist input bound assumptions)
  and strengthen the postconditions with abs(r) <= 6657.
- Add matching input/output bound CBMC contracts.

Also includes the earlier CBMC base-multiplication bound refinements
and the cassert()->mlk_assert() correction.

Signed-off-by: Dusan Kostic <dkostic@amazon.com>
@hanno-becker hanno-becker added benchmark this PR should be benchmarked in CI and removed benchmark this PR should be benchmarked in CI labels Aug 28, 2026

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 13433 cycles 13460 cycles 1.00
ML-KEM-512 encaps 15782 cycles 15785 cycles 1.00
ML-KEM-512 decaps 21311 cycles 21325 cycles 1.00
ML-KEM-768 keypair 22874 cycles 22835 cycles 1.00
ML-KEM-768 encaps 25295 cycles 25297 cycles 1.00
ML-KEM-768 decaps 33173 cycles 33118 cycles 1.00
ML-KEM-1024 keypair 33290 cycles 33220 cycles 1.00
ML-KEM-1024 encaps 36932 cycles 37031 cycles 1.00
ML-KEM-1024 decaps 47157 cycles 47147 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

@oqs-bot oqs-bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Graviton5 (no-opt)

Details
Benchmark suite Current: 2984fbe Previous: 0608277 Ratio
ML-KEM-512 keypair 27345 cycles 27360 cycles 1.00
ML-KEM-512 encaps 31709 cycles 31620 cycles 1.00
ML-KEM-512 decaps 40426 cycles 40530 cycles 1.00
ML-KEM-768 keypair 44011 cycles 44086 cycles 1.00
ML-KEM-768 encaps 50696 cycles 50589 cycles 1.00
ML-KEM-768 decaps 62149 cycles 62266 cycles 1.00
ML-KEM-1024 keypair 68288 cycles 68264 cycles 1.00
ML-KEM-1024 encaps 75988 cycles 76069 cycles 1.00
ML-KEM-1024 decaps 90665 cycles 90613 cycles 1.00

This comment was automatically generated by workflow using github-action-benchmark.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark this PR should be benchmarked in CI CBMC DO-NOT-MERGE enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants