Skip to content

unsloth: repin Inkling, #241 and #243 past b11453, drop #247, pin --moe-cache-mib auto (#251) - #252

Merged
danielhanchen merged 4 commits into
masterfrom
pin-moe-cache-auto-repin
Oct 8, 2026
Merged

danielhanchen merged 4 commits into
masterfrom
pin-moe-cache-auto-repin

Conversation

@danielhanchen

@danielhanchen danielhanchen commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Tonight's nightly would fail at Resolve tag. The base tag has moved past b11453, and three pins no longer merge onto it:

Pin Conflict Cause
ggml-org#25731 (TML Inkling) 11 files Upstream K2 Horizon (ggml-org#29535) took vocab pre-type id 61 and edits the same registration blocks; ggml-org#29442 changed the cuBLAS ldc path.
#243 (glm5-next compute buffer) src/models/glm5-next.cpp Upstream ggml-org#30042 removed the gather path of the glm5-next sparse attention.
#241 (carry) src/models/models.h Upstream ggml-org#29928 now ships the GLM5-Next MTP (NextN) graph itself.

One refused pin aborts the whole merge.

What changed here

Pin Old New
ggml-org#25731 Inkling efd2b13 fddee40
#241 carry b96a713 9dd7972
#243 glm5-next compute buffer aba4a4c f67cec2
#247 EmbeddingGemma-2 73c2f73 removed
#251 --moe-cache-mib auto 13f4c71 (new)

Details per pin:

Verification

  • Merge simulation. I ran the unsloth-prebuilt.yml merge loop (diff3, then additive_merge.py) on b11465 and on b11475. On both tags all 15 pins merge: 13 cleanly and 2 additively (Inkling and kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes #70), with no refusals.
  • pin_contract.py: all 15 pins are intact on both tags.
  • feature_matrix.py --gpu on the composed b11475 tree built with CUDA on B200: all 7 features demonstrated.
    • diffusion-gemma 3/3 device rows.
    • glm5next 192/192 cases.
    • inkling 13/13 cases plus the projector registry.
    • iq1-narrow-grids 3x 13/13.
    • kimi-k3, qwen4exp-mtp 3/3 rows and 4/4 cases, projector-registry.
  • Real models on the same composed tree:
    • Inkling-Small UD-IQ1_S: text and vision, through llama-mtmd-cli --jinja.
    • DiffusionGemma 26B Q4_K_M through llama-diffusion-cli: 60.9 tok/s.
    • gemma-4 E4B and 26B-A4B, Qwen3.6-35B-A3B: coherent output, also with --moe-cache-mib auto under a capped VRAM budget, with the same text.
    • Qwen3.8-Flash-Next with its MTP draft: draft acceptance 0.78 / 0.65.
    • GLM-5.3-Flash with --spec-type draft-mtp: works on both the legacy glm5next GGUF (via Load GLM-5-Next GGUFs converted with the earlier glm5next arch name #239) and the renamed glm5-next GGUF, with acceptance 0.65 / 0.61.
    • --moe-cache-mib auto with MTP: Qwen3.8 and GLM produce the same text and acceptance.
  • Carry what the dropped GLM-5-Next, Qwen MTP and readahead pins had beyond upstream #241 head on its own: test-llama-archs passes all 1319 cases.

Update 10-08: ggml-org#29887 and ggml-org#30112 merged upstream

ggml-org#29887 (the MoE expert cache) merged on 10-07, and ggml-org#30112 (the same cache over multiple GPUs) merged on 10-08 at c811cb8f0. With the pins as first pushed here, the merge loop on b11496 refused two of them:

Both branches now merge upstream master c811cb8f0:

Merge loop with pin_contract.py, as unsloth-prebuilt.yml runs it:

Base Merge pin_contract
b11496 15/15, 2 additive all 15 intact
b11500 15/15, 2 additive all 15 intact
c811cb8f0 (master, with ggml-org#30112) 15/15, 2 additive all 15 intact
b11480 15/15, 2 additive #251 flagged

The b11480 flag comes from one upstream chat-peg-parser line that Inkling's merge combines with its own change. The nightly takes the newest tag that is at least 6 h old, which is already b11500 or later, so it does not pick b11480 again.

Pin preflight harness fix

The first preflight run here (base b11461) merged all 15 pins and found them all intact, then reported glm5next as unproven: its only CPU row stopped right after the config column. That is a harness race, not the pins. test-llama-archs prints through common_log, whose worker thread is leaked at exit, so lines still queued when main returns can be lost. Locally, on one busy core, 5 of 60 runs of test-llama-archs -a glm5-next lost their tail the same way, and the full composed tree on b11461 passes every CPU probe.

feature_matrix.py now retries an arch probe up to 4 times while the Summary: line is missing, and otherwise fails with a message that names the truncation. Two new cases in test_feature_matrix.py cover it; both fail on the old script. 25 runs of the glm5next probe pinned to one core all pass.

The check_workflow_triggers.py failure (release-publish.yml uses workflow_run without the waiver comment) also fails on master since the b11436 sync and is not touched here.

Not covered here: ROCm, Metal and Windows builds. CI runs those.

…53, drop #247, pin #251

Upstream K2 Horizon (ggml-org#29535), the glm5-next gather removal (ggml-org#30042) and the
GLM5-Next MTP graph (ggml-org#29928) broke the Inkling, #243 and #241 pins. Each
branch now carries a merge of upstream master. #247 is in every tag from
b11452 on. #251 adds --moe-cache-mib auto on top of ggml-org#29887.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-07T15:44:03.046073Z 57df015 PR opened
🔒 Security Review ✅ Completed 2026-10-07T15:45:45.603801Z 57df015 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

test-llama-archs prints through common_log, whose worker thread is leaked at
exit, so lines still queued when main returns can be lost. The pin preflight
for this branch read a glm5-next row that stopped after its config column and
reported the feature as unproven. Locally, on one busy core, 5 of 60 runs of
test-llama-archs -a glm5-next lost their tail the same way.

Retry the probe up to 4 times while the Summary line is missing, and fail with
a message that names the truncation if it never arrives.
…gml-org#30112 merged

ggml-org#29887 (the MoE expert cache) merged on 10-07 and ggml-org#30112 (the cache over
multiple GPUs) on 10-08. On b11496 and later the old Inkling head refused in six
files and #251, which carried its own copy of ggml-org#29887, refused in eight.

Both branches now merge upstream master c811cb8. #251 keeps upstream's cache
lines intact, so the merged tree holds both pins in full.
#251 drops --moe-cache-static and the NextN/MTP cache filter, leaving the
auto sizing and the shared-MTP re-measure (+165/-17 against upstream).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant