Repository navigation
unsloth: repin Inkling, #241 and #243 past b11453, drop #247, pin --moe-cache-mib auto (#251) - #252
Merged
Merged
Conversation
…53, drop #247, pin #251 Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, #243 and #241 pins. Each branch now carries a merge of upstream master. #247 is in every tag from b11452 on. #251 adds --moe-cache-mib auto on top of ggml-org#29887.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
test-llama-archs prints through common_log, whose worker thread is leaked at exit, so lines still queued when main returns can be lost. The pin preflight for this branch read a glm5-next row that stopped after its config column and reported the feature as unproven. Locally, on one busy core, 5 of 60 runs of test-llama-archs -a glm5-next lost their tail the same way. Retry the probe up to 4 times while the Summary line is missing, and fail with a message that names the truncation if it never arrives.
…gml-org#30112 merged ggml-org#29887 (the MoE expert cache) merged on 10-07 and ggml-org#30112 (the cache over multiple GPUs) on 10-08. On b11496 and later the old Inkling head refused in six files and #251, which carried its own copy of ggml-org#29887, refused in eight. Both branches now merge upstream master c811cb8. #251 keeps upstream's cache lines intact, so the merged tree holds both pins in full.
#251 drops --moe-cache-static and the NextN/MTP cache filter, leaving the auto sizing and the shared-MTP re-measure (+165/-17 against upstream).
This was referenced Oct 8, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tonight's nightly would fail at Resolve tag. The base tag has moved past b11453, and three pins no longer merge onto it:
ldcpath.src/models/glm5-next.cppsrc/models/models.hOne refused pin aborts the whole merge.
What changed here
efd2b13fddee40b96a7139dd7972aba4a4cf67cec273c2f73--moe-cache-mib auto13f4c71(new)Details per pin:
tokenizer.ggml.preas a string.-ctk, and the macOS row prefetch.pin_contract.pyflagged it).74db68e), plus--moe-cache-mib auto, which sizes that cache from the fit. Its head is mainline + llama : add a GPU cache for MoE experts kept in host memory ggml-org/llama.cpp#29887 + three commits, so it merges independently of the other pins. It only acts when the flag is passed.feature-checks.jsonlists it underuncheckedwith the reason.Verification
unsloth-prebuilt.ymlmerge loop (diff3, thenadditive_merge.py) on b11465 and on b11475. On both tags all 15 pins merge: 13 cleanly and 2 additively (Inkling and kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes #70), with no refusals.pin_contract.py: all 15 pins are intact on both tags.feature_matrix.py --gpuon the composed b11475 tree built with CUDA on B200: all 7 features demonstrated.llama-mtmd-cli --jinja.llama-diffusion-cli: 60.9 tok/s.--moe-cache-mib autounder a capped VRAM budget, with the same text.--spec-type draft-mtp: works on both the legacyglm5nextGGUF (via Load GLM-5-Next GGUFs converted with the earlier glm5next arch name #239) and the renamedglm5-nextGGUF, with acceptance 0.65 / 0.61.--moe-cache-mib autowith MTP: Qwen3.8 and GLM produce the same text and acceptance.test-llama-archspasses all 1319 cases.Update 10-08: ggml-org#29887 and ggml-org#30112 merged upstream
ggml-org#29887 (the MoE expert cache) merged on 10-07, and ggml-org#30112 (the same cache over multiple GPUs) merged on 10-08 at
c811cb8f0. With the pins as first pushed here, the merge loop on b11496 refused two of them:Both branches now merge upstream master
c811cb8f0:fddee40: keeps both sides incommon/chat-peg-parser.cppandtools/mtmd/clip.cpp.13f4c71: upstream's cache plusautoonly, slimmed to +165/-17 in 9 files againstc811cb8f0. Upstream's cache files only gain the minimum-size helper.Merge loop with
pin_contract.py, asunsloth-prebuilt.ymlruns it:pin_contractc811cb8f0(master, with ggml-org#30112)The b11480 flag comes from one upstream
chat-peg-parserline that Inkling's merge combines with its own change. The nightly takes the newest tag that is at least 6 h old, which is already b11500 or later, so it does not pick b11480 again.Pin preflight harness fix
The first preflight run here (base b11461) merged all 15 pins and found them all intact, then reported glm5next as unproven: its only CPU row stopped right after the config column. That is a harness race, not the pins.
test-llama-archsprints throughcommon_log, whose worker thread is leaked at exit, so lines still queued whenmainreturns can be lost. Locally, on one busy core, 5 of 60 runs oftest-llama-archs -a glm5-nextlost their tail the same way, and the full composed tree on b11461 passes every CPU probe.feature_matrix.pynow retries an arch probe up to 4 times while theSummary:line is missing, and otherwise fails with a message that names the truncation. Two new cases intest_feature_matrix.pycover it; both fail on the old script. 25 runs of the glm5next probe pinned to one core all pass.The
check_workflow_triggers.pyfailure (release-publish.ymlusesworkflow_runwithout the waiver comment) also fails on master since the b11436 sync and is not touched here.Not covered here: ROCm, Metal and Windows builds. CI runs those.