Skip to content

feat(server): add ft info and a pre-load memory preflight - #595

Open
KarrAcaRn wants to merge 8 commits into
FlashML-org:mainfrom
KarrAcaRn:feat/ft-info
Open

KarrAcaRn wants to merge 8 commits into
FlashML-org:mainfrom
KarrAcaRn:feat/ft-info

Conversation

@KarrAcaRn

Copy link
Copy Markdown

I had an issue with a 21GB model not loading into a 24GB card.
It is possible to load this model but only with the right settings.
And i dont like that you only realise this after 10 minutes of loading and not at the begining.

It also makes it possible to have a better understanding of a model.

What

ft info <model> [ft serve flags] forecasts what ft serve would put on the GPU, and whether it fits, without reading a single weight. ft serve runs the same forecast between building the model and loading its weights, and refuses a configuration that cannot fit before spending minutes on the load.

Today a configuration that does not fit fails only after all weights are loaded, with an assert deep in KV sizing, and there is no way to see the trade-off between context length, concurrent requests and the MoE expert cache without trial and error.

How

  • The forecast builds the model on the meta device, exactly as Engine.__init__ does, so the resident weights are the tensors the loader would materialize.
  • Every pool is sized by the engine's own code, not a copy of it: _startup_kv_budget, the pool family's kv_cost / plan_num_pages, state_pool_bytes, and plan_moe_cache_auto. Two small refactors make these reusable: plan_num_pages is split out of solve_num_pages, and the moe-cache-auto plan is shared.
  • Only what the engine does not price up front is a heuristic (load overhead, prefill activations, CUDA graphs). The verdict never refuses on a heuristic alone, only when the engine's own solve would fail.
  • Free memory is read before the arch probe creates this process's CUDA context, so the context is not counted twice.
  • When the verdict is not "fits", ft info prints tips: each flag change is re-evaluated through the same forecast (--memory-ratio, --max-prefill-length, --kv-reserve-tokens, --moe-strategy, --cache-type, ...), plus a suggested combination.
  • ft serve preflight:
    • "fits" -> one info line.
    • "tight" -> a warning with the forecast and the tips.
    • "does not fit" -> refuses with the reason and a suggested fix. --skip-preflight loads anyway.
  • ft info --json gives the same report for tools. --gpu-memory-gib / --gpu-free-gib forecast for another GPU. Documented in docs/cli.md.

Tested

RTX 4090 24 GB (sm_89), driver 595.91.07, CUDA 13.1, 8-core CPU, 30 GB RAM, branch on main d3512b4.
Forecast vs. the engine's own log on a real start:

Checkpoint Command Forecast Actual
RadixArk/Qwen3.8-27B-NVFP4 ft info/serve <model> --text-model-only --max-running-requests 1 --memory-ratio 0.95 12931 KV tokens, 0.87 GiB free after init 12688 KV tokens, 0.86 GiB free
RedHatAI/Qwen3.6-35B-A3B-NVFP4 same flags 9755 expert slots, 8276 KV tokens 9859 expert slots, 8236 KV tokens
  • ft serve RadixArk/Qwen3.8-27B-NVFP4 --text-model-only --memory-ratio 0.5 is refused after ~15 s with "no room for the KV cache" and a suggested fix, before any weight is loaded.
  • pytest tests (-m "not slow"): new tests in tests/engine/test_forecast.py, tests/server/test_info.py and tests/kvcache/test_pool_sizing_surface.py pass. The 5 failures (test_fp8_pertensor_linear x2, qwen4_exp/test_qsa_backend chunked prefill x3) fail identically on main.

Follow-ups

The forecast prices exactly what main does today.
Once these open PRs are merged, small follow-up PRs extend it to match:

KarrAcaRn added 7 commits October 3, 2026 21:51
Assisted-by: Claude Opus 5.5
(cherry picked from commit 8dec696)
Assisted-by: Claude Opus 5.5
(cherry picked from commit 4a62081)
Assisted-by: Claude Opus 5.5
(cherry picked from commit 5962301)
…ts weights

Assisted-by: Claude Opus 5.5
(cherry picked from commit 056c6b8)
Assisted-by: Claude Opus 5.5
(cherry picked from commit 5126b3c)
…context

The context was counted twice: once in the reading and once in the 400 MiB allowance.

Assisted-by: Claude Opus 5.5
(cherry picked from commit 2425bde)
…am reserve

On upstream main the engine does not charge the attention workspace to the
pool budget and has no --vram-reserve-mb, so the forecast must not price
either: the workspace stays an item outside the budget (FlashInfer 256 MiB,
TRT-LLM 128 MiB, as allocated by attention/fi.py and attention/trtllm.py),
and the reserve step is gone. The sizing test no longer asserts that CUDA is
untouched (the arch probe creates the context on a GPU machine); instead a
test pins that free memory is read before that probe.

Assisted-by: Claude Opus 5.5
@KarrAcaRn

KarrAcaRn commented Oct 4, 2026 •

Copy link
Copy Markdown
Author

I'm think this PR would have given a better error message in the case of #561

GLM-4.x stores attention and embedding as DF11, whose buffers are sized by the
data at load and are empty placeholders in the meta build, so the forecast
counted 0 for them. GLM-4.5-Air: 3.25 -> 10.6 GiB on the GPU (measured 10.7-11.2
bits per weight on its tensors).

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant