Same Exam, Different Scores: A Tokenizer-Grounded Audit of Evaluation-Protocol Sensitivity on Chinese Benchmarks
NLPCC 2026 — Oral, Resources and Evaluation track. Kun Zhang · Chunwei Xia (corresponding) — School of Computing, University of Leeds
One sentence. The standard way of scoring multiple-choice LLM benchmarks is mathematically undefined for many Chinese models — and you can predict exactly which ones on a laptop, before spending a single GPU-hour.
一句话。 主流 MCQ 评测用的「首 token argmax」打分法,在许多中文模型上在数学上是未定义的 ——四个选项标签共享同一个首 token 时,argmax 比的是平局打破策略,不是模型能力。 哪些「模型 × 标签集」组合会退化,只由 tokenizer 决定,可在跑任何 GPU 之前用 CPU 预测。 我们预注册了这些预测,并在真实运行中 64/64 全部命中。
The point of an audit paper is that you should not have to take its word for it.
make smoke-test # ~5 s: materialize records + tokenizer audit, then check the headline story
make analysis # minutes: regenerate the degeneracy matrix over ~2.5M scoring records
make figures # regenerate the three headline figuresOr run the GPU-free audit directly — it downloads tokenizers only: no GPU, no model weights, no benchmark data.
pip install -r requirements-audit.txt
python src/cube/tokenizer_audit.pyRun this against your own label set before you trust a first-token score.
Multiple-choice evaluation silently fixes a protocol — how the answer label is scored, which label symbols are used, the shot count, the option order — that the benchmark leaves unspecified and that leaderboards set differently.
On Chinese benchmarks one of those choices can break the standard scorer outright. When a model's tokenizer maps all four answer labels onto a shared first token, first-token-argmax scoring is mathematically undefined: the first token carries no model-discriminative signal, so its output is a tie-break, not a score.
Which (model, label-set) pairs degenerate is a property of the tokenizer alone, so it is predictable before any model is run, on a laptop CPU. We pre-registered those predictions and confirmed all 64/64 label-scoring cells on the real runs; a full-label-sequence scorer recovers genuine accuracy in every flagged cell.
| Models | 18 checkpoints (16 label-scoring + 2 generation-only R1-Distills), 7 tokenizer families — Qwen2.5/Qwen3, GLM-4, InternLM3, Yi-1.5, Llama-3.1, Gemma-3, MAP-Neo |
| Benchmarks | C-Eval + CMMLU, 12,925 shared items |
| Label sets | ABCD · ABCD (full-width) · 甲乙丙丁 (Heavenly Stems) · ①②③④ (circled) |
| Degenerate cells | 12 of 64 (16 label-scoring models × 4 label sets) — every one predicted from the tokenizer before any GPU run |
| Rank instability | normalized Spearman footrule 0.472 between two defensible protocols (~77× the within-protocol noise floor's mean) under the strict pre-registered extractor; 0.167, still 6× the floor's 95th percentile, under a lenient re-reading of failed generations |
| Released records | ~2.5M scoring records across the protocol cube, archived under a citable DOI |
CPU, minutes GPU CPU
┌──────────────────────┐ ┌────────────────────────┐ ┌──────────────────┐
│ tokenizer audit │ │ scoring / generation │ │ analysis │
│ ──────────────── │ │ harness (vLLM) │ │ ───────── │
tokenizers→ predict which │──▶│ M1-seq · M1-strict │──▶│ pre-registered │─▶ figures
│ (model,label) cells │ │ M2–M5 cloze · M6 gen │ │ tests, degeneracy│ tables
│ degenerate │ │ 18 ckpts × 12,925 q │ │ matrix, footrule │ artifact
└──────────────────────┘ └────────────────────────┘ └──────────────────┘
│ │
pre-registered predictions ─────── confirmed 64/64 ─────────────┘
The design is predict-then-run: every claim that could be settled without a GPU was
frozen as a falsifiable prediction first, and the GPU runs were the test — not the
exploration. Prompt templates, label sets, model revisions and serving config are pinned
in config/; the pre-registration freeze commit is f2d4a63.
src/cube/
tokenizer_audit.py # the GPU-free audit: which cells degenerate, from the tokenizer alone
new_tokenizer_audit.py # re-runnable audit over additional checkpoints
harness/ # scoring + generation harness (vLLM); M1-seq, M1-strict, M2-M5 cloze, M6 generation
analysis/ # per-claim analyses (d6_*.py), pre-registered tests (stats_t*.py), figures (fig*.py)
release/ # artifact bundling, leakage audit, checksums, smoke test
config/ # frozen prompt templates, label sets, pinned model revisions, serving config
scripts/ # seeding and run supervision
The audit reports, for every (tokenizer family, label set) pair, whether the four labels' in-context first tokens are all distinct — i.e. whether first-token-argmax scoring is defined there at all. Expected output, which is the paper's pre-registered Table 2:
qwen2.5 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
qwen3 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
r1-distill-qwen {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'DEGEN'}
glm4 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
internlm3 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
yi1.5 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'DEGEN', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
map-neo {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
llama3.1 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max2tok', 'L4_circled': 'max2tok'}
gemma3 {'L1_halfwidth': 'max1tok', 'L2_fullwidth': 'max1tok', 'L3_ganzhi': 'max1tok', 'L4_circled': 'max1tok'}
DEGEN— the four labels' first tokens are not all distinct, so first-token-argmax is undefined. Eleven of the twelve are all-way collisions; InternLM3 atL2is a partial collision (B and D share a first token, A and C do not) and is conservatively scored the same way — seesrc/cube/analysis/partial_collision.py.maxNtok— the longest label in that cell. Note Llama-3.1'smax2tok: multi-token labels whose first tokens still differ are fine (fertility, not collision).- Llama-3.1 and Gemma-3 are gated on the Hugging Face Hub; without access those two lines
are reported as
[skip]on stderr and the rest of the audit still runs.
The audit loads tokenizers with
trust_remote_code=True, so for the two families that ship custom tokenizer code (InternLM3, MAP-NEO) it downloads and executes Python from those Hugging Face repos on your machine.
The per-item records, the frozen pre-registration, the deviation log and the statistics outputs are not in this repository — they are archived as a citable artifact:
https://doi.org/10.17605/OSF.IO/53C87
The artifact ships per-item records for all 18 models across the protocol cube (label
log-likelihoods, four-way option distributions, and generations), keyed by uid and
sha256(question)[:16] and carrying no benchmark text, so every protocol slice can be
re-scored without re-running a model and without redistributing C-Eval or CMMLU items. It
also contains the artifact checklist mapping each paper result to its reproduction path,
and the pre-registration freeze commit (f2d4a63) of the development repository.
Generations are scrubbed of option strings and any remaining ≥20-char question-text span,
and verified leak-free by src/cube/release/audit_leakage.py.
@inproceedings{zhang2026sameexam,
title = {Same Exam, Different Scores: A Tokenizer-Grounded Audit of
Evaluation-Protocol Sensitivity on {C}hinese Benchmarks},
author = {Zhang, Kun and Xia, Chunwei},
booktitle = {Proceedings of the CCF International Conference on Natural
Language Processing and Chinese Computing (NLPCC)},
year = {2026},
}Please also cite the artifact when you use the records: Zhang, K., & Xia, C. (2026). ZH-ProtocolCube. OSF. https://doi.org/10.17605/OSF.IO/53C87
Code in this repository is Apache-2.0 (see LICENSE). The released records and derived data in the OSF artifact are CC-BY-NC-SA-4.0. C-Eval and CMMLU remain under their own licences and are neither included here nor redistributed by the artifact.