Skip to content

unsloth: pin #253 in place of #241 to drop the CUDA graph cache cap (stacked on #252) - #254

Open
danielhanchen wants to merge 3 commits into
pin-moe-cache-auto-repinfrom
unsloth/pin-253-cuda-graph-cap
Open

danielhanchen wants to merge 3 commits into
pin-moe-cache-auto-repinfrom
unsloth/pin-253-cuda-graph-cap

Conversation

@danielhanchen

@danielhanchen danielhanchen commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Stacked on #252. Takes #252's pin set and puts #253 in #241's slot, so the nightly ships #241 without the 64 entry CUDA graph cache cap behind the --split-mode tensor decode regression in unslothai/unsloth#12468 (2.4x to 4.3x slower than the official ggml-org build on 2x RTX 5070 Ti, 1.57x on 2x B200). Against #252 this is a 2 line change.

Why #253 replaces #241 instead of following it

#253's head (20dd977ea) is #241's head 9dd797259, which #252 pins, merged with the one commit that removes the cap. Its tree is #241's with only ggml/src/ggml-cuda/common.cuh changed (+5 -14), so it carries everything #241's pin does.

Listing #253 after #241 cannot work: pin_contract.py checks that every line a pin adds is still in the merged tree, so #241 would fail on the removed max_cuda_graphs line, and unsloth-prebuilt.yml exits on that check. Taking #241's slot keeps the order and the composition, minus the cap.

feature-checks.json moves #241's unchecked reason to #253, noting that the cap only shows with --split-mode tensor on two GPUs.

Pin set merges additively

Replayed locally the way unsloth-prebuilt.yml does it (diff3 merge per pin, then scripts/unsloth/additive_merge.py on a conflict), onto b11491, the base the preflight picked today:

#252 this PR
pins that merge 15 of 15 15 of 15
resolved by additive_merge.py ggml-org#25731, #70 ggml-org#25731, #70
merge_checks.py clean clean
pin_contract.py all 15 intact all 15 intact
max_cuda_graphs lines in the merged common.cuh 2 0
shape keyed ggml_cuda_graph_get_key present present
  • The merged tree passes the CPU compile gate (llama and mtmd) and builds llama-bench with CUDA (sm_100).
  • On the merged tree's own llama-bench (Qwen3-4B Q4_K_M, -sm tensor, 2x B200, shared GPUs on a loaded host, so this shows engagement rather than a speed figure), CUDA graphs beat GGML_CUDA_DISABLE_GRAPHS=1 in all 3 rounds: 150 vs 85, 194 vs 30, 93 vs 50 tok/s. With the cap the two were equal.
  • The pin coverage lint passes (15 pins, 7 checked, 8 knowingly unchecked).
  • The replay was checked against a known result first: the pin set of the b11443-mix-d65395f release, replayed onto b11443, merges all 15 pins (Add TML Inkling architecture ggml-org/llama.cpp#25731, kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes #70 and EmbeddingGemma-2 support #247 additively), as that nightly did.

The preflight run on this branch before it was stacked failed at ggml-org#25731 on b11491. That is the Inkling conflict #252 fixes, and it never reached #253.

Order

Merge #252 first, then this. If #252 changes again, this needs only the #241 line re-swapped to a #253 head that contains #241's new head.

#253 is #241's pinned commit plus one commit that removes the 64 entry
cap on the CUDA graph cache, which evicts graphs still in use under
--split-mode tensor and makes decode as slow as running without CUDA
graphs (unslothai/unsloth#12468). Upstream has no count cap.

It replaces #241 rather than following it: a later pin that removes
lines an earlier pin adds fails pin_contract.py by design, and #253
already carries everything #241's pin does.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-08T11:43:45.867573Z 1654d1f PR opened
🔒 Security Review ✅ Completed 2026-10-08T11:45:39.241135Z 1654d1f PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

…-cap

# Conflicts:
#	scripts/unsloth/feature-checks.json
#	scripts/unsloth/pr-set.json
@danielhanchen
danielhanchen changed the base branch from master to pin-moe-cache-auto-repin October 8, 2026 16:25
@danielhanchen danielhanchen changed the title unsloth: pin #253 in place of #241 to drop the CUDA graph cache cap unsloth: pin #253 in place of #241 to drop the CUDA graph cache cap (stacked on #252) Oct 8, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant