You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
EXPERIMENTS.md E5: control = run-006-r7-muon verbatim, single delta =
a 3-step multi-token-prediction head (~1.1M params, 0.4%) on the
student decoder, aux CE at beta 0.15, head discarded at inference so
the shipped artifact stays vanilla and size-identical. Source: Hy4's
native MTP mapped as a TRAINING auxiliary (TODO 07-hy4) — serving-side
speculation stays parked, decode measured non-binding.
gpu/mtp.py (MTPHead + build/mtp_named), distill_sequence wiring
(head before optimizer so Muon sees its matrices; mtp_head.pt saved
beside student.pt for resume; provenance-only copy under best/),
spec ara-diac-small-2-mtp -> run-007-r7-muon-mtp. 3 unit tests
(shift/mask semantics, uniform-logit baseline, naming).
Gate (pre-agreed): adopt at <=4.5218 full-set windowed DER; honest
report in [4.5218, 4.8218). Prediction: 4.5-4.75.
0 commit comments