Fix quantize-lm on macOS (#21), gradient checkpointing freezing blocks, RoPE grader accepting no-ops - #31
Open
Tar-ive wants to merge 3 commits into
Open
Fix quantize-lm on macOS (#21), gradient checkpointing freezing blocks, RoPE grader accepting no-ops#31Tar-ive wants to merge 3 commits into
Tar-ive wants to merge 3 commits into
Conversation
…ining loop (Exorust#21) - Pick a quantized engine when the default is 'none' (Apple Silicon), the same guard the grader already uses. Previously quantize_dynamic raised "Didn't find engine for operation quantized::linear_prepack NoQEngine". - Save the float state_dict and re-quantize after loading. A dynamically quantized LSTM state_dict holds torch.ScriptObject packed params, which torch.load(weights_only=True) - the default since 2.6 - refuses to unpickle. Dynamic quantization is deterministic, so the reloaded model is identical. - The model ends in a softmax, so train with NLLLoss on log-probabilities instead of CrossEntropyLoss (a second softmax), which pinned the loss at ~ln(50): 3.9118 -> 3.9097 over 5 epochs, now 3.9032 -> 3.8048. - Import quantize_dynamic from torch.ao.quantization (torch.quantization is the deprecated alias). - Solution: add the float-vs-int8 comparison the problem statement asks for (weight size and top-1 agreement). Clear outputs produced by the old code.
autograd only records a custom Function, and later calls its backward, when at least one tensor argument requires grad. The block parameters fn closes over are not arguments, so with plain input data (the notebook's own training test, and the first layer of any real network) backward never ran: only the head trained and all 96 block parameters had grad None, yet the test passed because the head alone brought the loss down. - Solution: checkpoint() passes a dummy requires_grad tensor when no argument requires grad; backward skips outputs that need no grad. - Training test now asserts every parameter receives a gradient with plain input (the previous solution fails it). - Question: hint and TODOs no longer steer to autograd.grad w.r.t. the inputs (which never reaches fn's parameters), and the validation treats a missing gradient as a failure instead of skipping it, matching the solution. - Grader: a checkpoint whose wrapped module gets no gradient now fails instead of being reported as "not verified" (a pass); new check for plain input. Hints updated and regenerated into problem.toml, problems.json and the MCP server's hint table.
Every existing value check was satisfied by returning (q, k) unchanged: the zero-angle check expects exactly that, and a no-op trivially preserves norms. Add check_quarter_turn: at cos=0, sin=1 the result must equal the solver's own rotate_half(q) and rotate_half(k), which also catches forgetting to rotate k. test_grader.py gains wrong-answer cases for RoPE (no-op, query-only) and gradient checkpointing (no dummy input, autograd.grad on inputs only) so these checks cannot regress to accepting broken code.
|
@Tar-ive is attempting to deploy a commit to the exorust's projects Team on Vercel. A member of the Team first needs to authorize it. |
|
Check out this pull request on See visual diffs & provide feedback on Jupyter Notebooks. Powered by ReviewNB |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #21. Also fixes gradient checkpointing silently leaving every checkpointed block untrained, and tightens the RoPE grader so a no-op implementation no longer passes.
1. quantize-lm fails on macOS / torch ≥ 2.6 (#21)
Reproduced on an M-series Mac with torch 2.9.0 (engine defaults to
'none', supported:['qnnpack']), in two stages exactly as reported:quantize_dynamic→RuntimeError: Didn't find engine for operation quantized::linear_prepack NoQEnginetorch.load(...)→UnpicklingError ... Unsupported global: GLOBAL torch.ScriptObject. A dynamically quantized LSTMstate_dictstores packed params astorch.ScriptObject, whichweights_only=True(the default since 2.6) rejects.Fix (question and solution notebooks share this scaffold, so both are updated):
'none'. This is the same guard the grader already has inquantize_lm.py.weights_only=Truepath. The alternative,weights_only=Falseor allow-listingtorch.ScriptObject, would teach an unsafe loading habit.from torch.ao.quantization import quantize_dynamic(torch.quantizationis the deprecated alias).Separate bug in the same notebook: the model ends in
nn.Softmaxand was trained withnn.CrossEntropyLoss, which applies a second softmax. That pinned the loss at ≈ ln(50); the committed output showed3.9118 → 3.9097over 5 epochs. The grader requires probabilities out of the model, so the fix keeps the softmax and trains withNLLLossonlog(probs). The loss now goes3.9032 → 3.8048.The solution also gains the comparison the problem statement asks for ("evaluate the quantized model's performance compared to the original model"): on this machine, weights go 0.97 MB → 0.27 MB with 100% top-1 agreement. Outputs produced by the old code were cleared.
2. Gradient checkpointing: checkpointed blocks never train
autograd only records a custom
Function, and only calls itsbackward, if at least one tensor argument requires grad. The block's parameters are closed over byfn, not passed as arguments. So when the input is plain data, which is the notebook's own training test and the first layer of any real network,backwardnever ran. Onlyheadtrained, all 96 block parameters hadgrad is None, and the test still passed because the head alone lowered the loss.checkpoint()passes a dummyrequires_grad=Truetensor when no argument requires grad.backwardskips outputs that need no grad (e.g. a frozenfn).96 parameters got no gradient.torch.autograd.gradw.r.t. the inputs, which never reachesfn's parameters. The validation now treats a missing gradient as a failure instead of skipping it, matching the solution.check_module_parameter_gradientsused to raiseSkipwhen the wrapped module got no gradient. The runner counts that as a pass, so the broken formulation was accepted. It now fails. A new check,check_plain_input_still_trains_parameters, covers the no-grad input case. The hints were corrected and regenerated intoproblem.toml/problems.json(generate.py hints+json). They were also synced into the MCP server'shints.generated.ts, which no script regenerates today.3. RoPE grader accepted a no-op
check("rotary-positional-embedding", rotate_half, apply)passed withapply = lambda q, k, cos, sin: (q, k). The zero-angle check expects exactly that, and a no-op trivially preserves norms. It also passed when onlyqwas rotated. The newcheck_quarter_turnchecks that atcos=0, sin=1the output equals the solver's ownrotate_half(q)androtate_half(k), so any consistent pairing convention still passes.Tests
All run on torch 2.9.0, macOS arm64, CPU:
python/test_grader.py: 14/14. This includes 5 new wrong-answer cases (RoPE no-op, RoPE query-only, checkpoint without dummy input, checkpoint withautograd.gradon inputs only, plus the correct versions), so these checks can't regress to accepting broken code.scripts/validate_specs.py rotary-positional-embedding gradient-checkpointing quantize-lm: all 3 repo solutions pass their updated specs.generate.py check: 68 manifests, 0 problems. The MCP server typechecks (tsc --noEmit).Notes / follow-ups (not in this PR)
torch.ao.quantizationemits a deprecation warning announcing its removal. Longer term this exercise should move totorchao(quantize_with an int8 dynamic config).batch_first) looks already fixed onmain(60d933e replacednn.TransformerEncoderLayerwith a hand-written batch-first layer; the current solution reaches ~98.5%) and could be closed.