Skip to content

Repository files navigation

ZH-ProtocolCube

Same Exam, Different Scores: A Tokenizer-Grounded Audit of Evaluation-Protocol Sensitivity on Chinese Benchmarks

NLPCC 2026 — Oral, Resources and Evaluation track. Kun Zhang · Chunwei Xia (corresponding) — School of Computing, University of Leeds

Artifact DOI Code license Data license Pre-registered

One sentence. The standard way of scoring multiple-choice LLM benchmarks is mathematically undefined for many Chinese models — and you can predict exactly which ones on a laptop, before spending a single GPU-hour.

一句话。 主流 MCQ 评测用的「首 token argmax」打分法,在许多中文模型上在数学上是未定义的 ——四个选项标签共享同一个首 token 时,argmax 比的是平局打破策略,不是模型能力。 哪些「模型 × 标签集」组合会退化,只由 tokenizer 决定,可在跑任何 GPU 之前用 CPU 预测。 我们预注册了这些预测,并在真实运行中 64/64 全部命中。


Verify the headline claim in ~5 seconds

The point of an audit paper is that you should not have to take its word for it.

make smoke-test     # ~5 s: materialize records + tokenizer audit, then check the headline story
make analysis       # minutes: regenerate the degeneracy matrix over ~2.5M scoring records
make figures        # regenerate the three headline figures

Or run the GPU-free audit directly — it downloads tokenizers only: no GPU, no model weights, no benchmark data.

pip install -r requirements-audit.txt
python src/cube/tokenizer_audit.py

Run this against your own label set before you trust a first-token score.


The finding

Multiple-choice evaluation silently fixes a protocol — how the answer label is scored, which label symbols are used, the shot count, the option order — that the benchmark leaves unspecified and that leaderboards set differently.

On Chinese benchmarks one of those choices can break the standard scorer outright. When a model's tokenizer maps all four answer labels onto a shared first token, first-token-argmax scoring is mathematically undefined: the first token carries no model-discriminative signal, so its output is a tie-break, not a score.

Which (model, label-set) pairs degenerate is a property of the tokenizer alone, so it is predictable before any model is run, on a laptop CPU. We pre-registered those predictions and confirmed all 64/64 label-scoring cells on the real runs; a full-label-sequence scorer recovers genuine accuracy in every flagged cell.

Models 18 checkpoints (16 label-scoring + 2 generation-only R1-Distills), 7 tokenizer families — Qwen2.5/Qwen3, GLM-4, InternLM3, Yi-1.5, Llama-3.1, Gemma-3, MAP-Neo
Benchmarks C-Eval + CMMLU, 12,925 shared items
Label sets ABCD · ABCD (full-width) · 甲乙丙丁 (Heavenly Stems) · ①②③④ (circled)
Degenerate cells 12 of 64 (16 label-scoring models × 4 label sets) — every one predicted from the tokenizer before any GPU run
Rank instability normalized Spearman footrule 0.472 between two defensible protocols (~77× the within-protocol noise floor's mean) under the strict pre-registered extractor; 0.167, still 6× the floor's 95th percentile, under a lenient re-reading of failed generations
Released records ~2.5M scoring records across the protocol cube, archived under a citable DOI

How it was built

             CPU, minutes                      GPU                       CPU
        ┌──────────────────────┐   ┌────────────────────────┐   ┌──────────────────┐
        │  tokenizer audit     │   │  scoring / generation  │   │  analysis        │
        │  ────────────────    │   │  harness (vLLM)        │   │  ─────────       │
tokenizers→ predict which     │──▶│  M1-seq · M1-strict    │──▶│ pre-registered   │─▶ figures
        │  (model,label) cells │   │  M2–M5 cloze · M6 gen  │   │ tests, degeneracy│   tables
        │  degenerate          │   │  18 ckpts × 12,925 q   │   │ matrix, footrule │   artifact
        └──────────────────────┘   └────────────────────────┘   └──────────────────┘
                  │                                                       │
          pre-registered predictions ─────── confirmed 64/64 ─────────────┘

The design is predict-then-run: every claim that could be settled without a GPU was frozen as a falsifiable prediction first, and the GPU runs were the test — not the exploration. Prompt templates, label sets, model revisions and serving config are pinned in config/; the pre-registration freeze commit is f2d4a63.

Repository layout

src/cube/
  tokenizer_audit.py         # the GPU-free audit: which cells degenerate, from the tokenizer alone
  new_tokenizer_audit.py     # re-runnable audit over additional checkpoints
  harness/                   # scoring + generation harness (vLLM); M1-seq, M1-strict, M2-M5 cloze, M6 generation
  analysis/                  # per-claim analyses (d6_*.py), pre-registered tests (stats_t*.py), figures (fig*.py)
  release/                   # artifact bundling, leakage audit, checksums, smoke test
config/                      # frozen prompt templates, label sets, pinned model revisions, serving config
scripts/                     # seeding and run supervision

Reading the audit output

The audit reports, for every (tokenizer family, label set) pair, whether the four labels' in-context first tokens are all distinct — i.e. whether first-token-argmax scoring is defined there at all. Expected output, which is the paper's pre-registered Table 2:

qwen2.5          {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
qwen3            {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
r1-distill-qwen  {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
glm4             {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN',   'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
internlm3        {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN',   'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
yi1.5            {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN',   'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
map-neo          {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
llama3.1         {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max2tok', 'L4_circled': 'max2tok'}
gemma3           {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
  • DEGEN — the four labels' first tokens are not all distinct, so first-token-argmax is undefined. Eleven of the twelve are all-way collisions; InternLM3 at L2 is a partial collision (B and D share a first token, A and C do not) and is conservatively scored the same way — see src/cube/analysis/partial_collision.py.
  • maxNtok — the longest label in that cell. Note Llama-3.1's max2tok: multi-token labels whose first tokens still differ are fine (fertility, not collision).
  • Llama-3.1 and Gemma-3 are gated on the Hugging Face Hub; without access those two lines are reported as [skip] on stderr and the rest of the audit still runs.

The audit loads tokenizers with trust_remote_code=True, so for the two families that ship custom tokenizer code (InternLM3, MAP-NEO) it downloads and executes Python from those Hugging Face repos on your machine.


Data and full reproduction

The per-item records, the frozen pre-registration, the deviation log and the statistics outputs are not in this repository — they are archived as a citable artifact:

https://doi.org/10.17605/OSF.IO/53C87

The artifact ships per-item records for all 18 models across the protocol cube (label log-likelihoods, four-way option distributions, and generations), keyed by uid and sha256(question)[:16] and carrying no benchmark text, so every protocol slice can be re-scored without re-running a model and without redistributing C-Eval or CMMLU items. It also contains the artifact checklist mapping each paper result to its reproduction path, and the pre-registration freeze commit (f2d4a63) of the development repository.

Generations are scrubbed of option strings and any remaining ≥20-char question-text span, and verified leak-free by src/cube/release/audit_leakage.py.


Citation

@inproceedings{zhang2026sameexam,
  title     = {Same Exam, Different Scores: A Tokenizer-Grounded Audit of
               Evaluation-Protocol Sensitivity on {C}hinese Benchmarks},
  author    = {Zhang, Kun and Xia, Chunwei},
  booktitle = {Proceedings of the CCF International Conference on Natural
               Language Processing and Chinese Computing (NLPCC)},
  year      = {2026},
}

Please also cite the artifact when you use the records: Zhang, K., & Xia, C. (2026). ZH-ProtocolCube. OSF. https://doi.org/10.17605/OSF.IO/53C87

Licensing

Code in this repository is Apache-2.0 (see LICENSE). The released records and derived data in the OSF artifact are CC-BY-NC-SA-4.0. C-Eval and CMMLU remain under their own licences and are neither included here nor redistributed by the artifact.

About

Tokenizer-grounded audit of evaluation-protocol sensitivity for multiple-choice LLM evaluation on Chinese benchmarks (NLPCC 2026)

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages