Repository navigation
Conversation
added 7 commits
October 3, 2026 21:51
Assisted-by: Claude Opus 5.5 (cherry picked from commit 8dec696)
Assisted-by: Claude Opus 5.5 (cherry picked from commit 4a62081)
Assisted-by: Claude Opus 5.5 (cherry picked from commit 5962301)
…ts weights Assisted-by: Claude Opus 5.5 (cherry picked from commit 056c6b8)
Assisted-by: Claude Opus 5.5 (cherry picked from commit 5126b3c)
…context The context was counted twice: once in the reading and once in the 400 MiB allowance. Assisted-by: Claude Opus 5.5 (cherry picked from commit 2425bde)
…am reserve On upstream main the engine does not charge the attention workspace to the pool budget and has no --vram-reserve-mb, so the forecast must not price either: the workspace stays an item outside the budget (FlashInfer 256 MiB, TRT-LLM 128 MiB, as allocated by attention/fi.py and attention/trtllm.py), and the reserve step is gone. The sizing test no longer asserts that CUDA is untouched (the arch probe creates the context on a GPU machine); instead a test pins that free memory is read before that probe. Assisted-by: Claude Opus 5.5
Author
|
I'm think this PR would have given a better error message in the case of #561 |
GLM-4.x stores attention and embedding as DF11, whose buffers are sized by the data at load and are empty placeholders in the meta build, so the forecast counted 0 for them. GLM-4.5-Air: 3.25 -> 10.6 GiB on the GPU (measured 10.7-11.2 bits per weight on its tensors). Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
I had an issue with a 21GB model not loading into a 24GB card.
It is possible to load this model but only with the right settings.
And i dont like that you only realise this after 10 minutes of loading and not at the begining.
It also makes it possible to have a better understanding of a model.
What
ft info <model> [ft serve flags]forecasts whatft servewould put on the GPU, and whether it fits, without reading a single weight.ft serveruns the same forecast between building the model and loading its weights, and refuses a configuration that cannot fit before spending minutes on the load.Today a configuration that does not fit fails only after all weights are loaded, with an assert deep in KV sizing, and there is no way to see the trade-off between context length, concurrent requests and the MoE expert cache without trial and error.
How
Engine.__init__does, so the resident weights are the tensors the loader would materialize._startup_kv_budget, the pool family'skv_cost/plan_num_pages,state_pool_bytes, andplan_moe_cache_auto. Two small refactors make these reusable:plan_num_pagesis split out ofsolve_num_pages, and the moe-cache-auto plan is shared.ft infoprints tips: each flag change is re-evaluated through the same forecast (--memory-ratio,--max-prefill-length,--kv-reserve-tokens,--moe-strategy,--cache-type, ...), plus a suggested combination.ft servepreflight:--skip-preflightloads anyway.ft info --jsongives the same report for tools.--gpu-memory-gib/--gpu-free-gibforecast for another GPU. Documented indocs/cli.md.Tested
RTX 4090 24 GB (sm_89), driver 595.91.07, CUDA 13.1, 8-core CPU, 30 GB RAM, branch on
maind3512b4.Forecast vs. the engine's own log on a real start:
RadixArk/Qwen3.8-27B-NVFP4ft info/serve <model> --text-model-only --max-running-requests 1 --memory-ratio 0.95RedHatAI/Qwen3.6-35B-A3B-NVFP4ft serve RadixArk/Qwen3.8-27B-NVFP4 --text-model-only --memory-ratio 0.5is refused after ~15 s with "no room for the KV cache" and a suggested fix, before any weight is loaded.pytest tests(-m "not slow"): new tests intests/engine/test_forecast.py,tests/server/test_info.pyandtests/kvcache/test_pool_sizing_surface.pypass. The 5 failures (test_fp8_pertensor_linearx2,qwen4_exp/test_qsa_backendchunked prefill x3) fail identically onmain.Follow-ups
The forecast prices exactly what
maindoes today.Once these open PRs are merged, small follow-up PRs extend it to match:
plan_moe_cache_auto. Until then it is priced only as memory outside the budget.--vram-reserve-mb): price the reserve and the expert-cache shrink it triggers.--kv-cache-dtype fp8/nvfp4): add KV-dtype tips showing the extra context each dtype buys.