Skip to content

Fix quantize-lm on macOS (#21), gradient checkpointing freezing blocks, RoPE grader accepting no-ops - #31

Open
Tar-ive wants to merge 3 commits into
Exorust:mainfrom
Tar-ive:fix/quantize-lm-macos-gradckpt-rope
Open

Tar-ive wants to merge 3 commits into
Exorust:mainfrom
Tar-ive:fix/quantize-lm-macos-gradckpt-rope

Conversation

@Tar-ive

@Tar-ive Tar-ive commented Sep 24, 2026

Copy link
Copy Markdown

Fixes #21. Also fixes gradient checkpointing silently leaving every checkpointed block untrained, and tightens the RoPE grader so a no-op implementation no longer passes.

1. quantize-lm fails on macOS / torch ≥ 2.6 (#21)

Reproduced on an M-series Mac with torch 2.9.0 (engine defaults to 'none', supported: ['qnnpack']), in two stages exactly as reported:

  • quantize_dynamic → RuntimeError: Didn't find engine for operation quantized::linear_prepack NoQEngine
  • after setting the engine, torch.load(...) → UnpicklingError ... Unsupported global: GLOBAL torch.ScriptObject. A dynamically quantized LSTM state_dict stores packed params as torch.ScriptObject, which weights_only=True (the default since 2.6) rejects.

Fix (question and solution notebooks share this scaffold, so both are updated):

  • Select a supported quantized engine when it is 'none'. This is the same guard the grader already has in quantize_lm.py.
  • Save the float weights and quantize again after loading. Dynamic quantization is deterministic, so the reloaded model is identical, and the load stays on the safe weights_only=True path. The alternative, weights_only=False or allow-listing torch.ScriptObject, would teach an unsafe loading habit.
  • from torch.ao.quantization import quantize_dynamic (torch.quantization is the deprecated alias).

Separate bug in the same notebook: the model ends in nn.Softmax and was trained with nn.CrossEntropyLoss, which applies a second softmax. That pinned the loss at ≈ ln(50); the committed output showed 3.9118 → 3.9097 over 5 epochs. The grader requires probabilities out of the model, so the fix keeps the softmax and trains with NLLLoss on log(probs). The loss now goes 3.9032 → 3.8048.

The solution also gains the comparison the problem statement asks for ("evaluate the quantized model's performance compared to the original model"): on this machine, weights go 0.97 MB → 0.27 MB with 100% top-1 agreement. Outputs produced by the old code were cleared.

2. Gradient checkpointing: checkpointed blocks never train

autograd only records a custom Function, and only calls its backward, if at least one tensor argument requires grad. The block's parameters are closed over by fn, not passed as arguments. So when the input is plain data, which is the notebook's own training test and the first layer of any real network, backward never ran. Only head trained, all 96 block parameters had grad is None, and the test still passed because the head alone lowered the loss.

  • Solution: checkpoint() passes a dummy requires_grad=True tensor when no argument requires grad. backward skips outputs that need no grad (e.g. a frozen fn).
  • Training test: asserts every parameter receives a gradient with plain input. The previous solution fails it with 96 parameters got no gradient.
  • Question notebook: the hint and TODOs no longer steer to torch.autograd.grad w.r.t. the inputs, which never reaches fn's parameters. The validation now treats a missing gradient as a failure instead of skipping it, matching the solution.
  • Grader: check_module_parameter_gradients used to raise Skip when the wrapped module got no gradient. The runner counts that as a pass, so the broken formulation was accepted. It now fails. A new check, check_plain_input_still_trains_parameters, covers the no-grad input case. The hints were corrected and regenerated into problem.toml / problems.json (generate.py hints + json). They were also synced into the MCP server's hints.generated.ts, which no script regenerates today.

3. RoPE grader accepted a no-op

check("rotary-positional-embedding", rotate_half, apply) passed with apply = lambda q, k, cos, sin: (q, k). The zero-angle check expects exactly that, and a no-op trivially preserves norms. It also passed when only q was rotated. The new check_quarter_turn checks that at cos=0, sin=1 the output equals the solver's own rotate_half(q) and rotate_half(k), so any consistent pairing convention still passes.

Tests

All run on torch 2.9.0, macOS arm64, CPU:

  • python/test_grader.py: 14/14. This includes 5 new wrong-answer cases (RoPE no-op, RoPE query-only, checkpoint without dummy input, checkpoint with autograd.grad on inputs only, plus the correct versions), so these checks can't regress to accepting broken code.
  • scripts/validate_specs.py rotary-positional-embedding gradient-checkpointing quantize-lm: all 3 repo solutions pass their updated specs.
  • Both fixed solution notebooks run end to end. The quantize-lm question scaffold also runs end to end with the solution's model.
  • generate.py check: 68 manifests, 0 problems. The MCP server typechecks (tsc --noEmit).

Notes / follow-ups (not in this PR)

  • torch.ao.quantization emits a deprecation warning announcing its removal. Longer term this exercise should move to torchao (quantize_ with an int8 dynamic config).
  • The PyPI package is still 0.1.0, so the grader changes reach users only after a release.
  • Issue In h3, we should pass batch_first=True to TransformerEncoderLayer #5 (transformer batch_first) looks already fixed on main (60d933e replaced nn.TransformerEncoderLayer with a hand-written batch-first layer; the current solution reaches ~98.5%) and could be closed.

…ining loop (Exorust#21)

- Pick a quantized engine when the default is 'none' (Apple Silicon), the same
  guard the grader already uses. Previously quantize_dynamic raised
  "Didn't find engine for operation quantized::linear_prepack NoQEngine".
- Save the float state_dict and re-quantize after loading. A dynamically
  quantized LSTM state_dict holds torch.ScriptObject packed params, which
  torch.load(weights_only=True) - the default since 2.6 - refuses to unpickle.
  Dynamic quantization is deterministic, so the reloaded model is identical.
- The model ends in a softmax, so train with NLLLoss on log-probabilities
  instead of CrossEntropyLoss (a second softmax), which pinned the loss at
  ~ln(50): 3.9118 -> 3.9097 over 5 epochs, now 3.9032 -> 3.8048.
- Import quantize_dynamic from torch.ao.quantization (torch.quantization is the
  deprecated alias).
- Solution: add the float-vs-int8 comparison the problem statement asks for
  (weight size and top-1 agreement). Clear outputs produced by the old code.
autograd only records a custom Function, and later calls its backward, when at
least one tensor argument requires grad. The block parameters fn closes over
are not arguments, so with plain input data (the notebook's own training test,
and the first layer of any real network) backward never ran: only the head
trained and all 96 block parameters had grad None, yet the test passed because
the head alone brought the loss down.

- Solution: checkpoint() passes a dummy requires_grad tensor when no argument
  requires grad; backward skips outputs that need no grad.
- Training test now asserts every parameter receives a gradient with plain
  input (the previous solution fails it).
- Question: hint and TODOs no longer steer to autograd.grad w.r.t. the inputs
  (which never reaches fn's parameters), and the validation treats a missing
  gradient as a failure instead of skipping it, matching the solution.
- Grader: a checkpoint whose wrapped module gets no gradient now fails instead
  of being reported as "not verified" (a pass); new check for plain input.
  Hints updated and regenerated into problem.toml, problems.json and the MCP
  server's hint table.
Every existing value check was satisfied by returning (q, k) unchanged: the
zero-angle check expects exactly that, and a no-op trivially preserves norms.
Add check_quarter_turn: at cos=0, sin=1 the result must equal the solver's own
rotate_half(q) and rotate_half(k), which also catches forgetting to rotate k.

test_grader.py gains wrong-answer cases for RoPE (no-op, query-only) and
gradient checkpointing (no dummy input, autograd.grad on inputs only) so these
checks cannot regress to accepting broken code.
@Tar-ive
Tar-ive requested a review from Exorust as a code owner September 24, 2026 18:49
@vercel

vercel Bot commented Sep 24, 2026

Copy link
Copy Markdown

@Tar-ive is attempting to deploy a commit to the exorust's projects Team on Vercel.

A member of the Team first needs to authorize it.

@review-notebook-app

Copy link
Copy Markdown

Check out this pull request on  ReviewNB

See visual diffs & provide feedback on Jupyter Notebooks.


Powered by ReviewNB

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Solution for quantize-lm does not work on macOS

1 participant