You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit ccb0be0
Browse filesBrowse the repository at this point in the historyBrowse files
Ronald Tse
committed
fix(imf): fp16 via torch-native half export — ORT converter broken on real ByT5
The onnxruntime float16 converter produced all-zero encoder hiddens on
the real khm-latn checkpoint (CER 1939pp, every sample mismatched) while
looking fine on the tiny fixture. Exporting the model under .half() is
exact on gold pairs; graph IO becomes float16 (int64 ids unchanged) and
_zero_pasts follows session dtypes.
Measured on 300 khm test samples: fp16 delta 0.43pp, int8 0.84pp —
quantization noise (argmax flips cascading under greedy decode), not
breakage. Whether lossy precisions may exceed the 0.2pp export-fidelity
bar is a policy call pending; fp32 measures 0.0pp.
|`precision`| enum |`fp32`\|`fp16`\|`int8`(fp16 = torch-native half export: float16 graph IO, int64 ids unchanged; runtimes read dtypes from the session — the ORT float16 converter produces all-zero hiddens on real ByT5 and must not be used) |
51
51
|`license`| str | non-empty (strict gate) |
52
52
|`trained_from`| str | repo + run/checkpoint id |
53
53
|`metrics`| list |`{name, value, protocol, source}`; `source` must be a `RESULTS.md#anchor` (strict gate) |
0 commit comments