Skip to content

fix(engine): apply shrinking cache resizes before growing ones in live rebuild (Closes #643) - #650

Open
void-mckenzie wants to merge 1 commit into
FlashML-org:mainfrom
void-mckenzie:fix/cache-rebuild-ordering
Open

void-mckenzie wants to merge 1 commit into
FlashML-org:mainfrom
void-mckenzie:fix/cache-rebuild-ordering

Conversation

@void-mckenzie

Copy link
Copy Markdown
Contributor

Closes #643.

Engine.rebuild_runtime_cache applied its three resizes in a fixed MoE -> KV -> mamba order
regardless of direction, and each pool frees only its own old tensors, so a mixed-direction
resize allocated the larger MoE cache while the old KV/mamba pools were still resident.
validate_rebuild prices only the target geometry, so a plan that passed the fit check could
OOM mid-rebuild (and the failed rollback after such an OOM is #526).

The rebuild now compares each pool family's old-vs-new footprint (new
BaseKVCachePool.rebuild_footprint classmethod reusing the kv_cost plumbing
validate_rebuild already uses; DSV4 overrides it with _dsv4_pool_sizes +
dsv4_pool_bytes, since its kv_cost takes no num_swa_pages) and runs the shrinking
resizes before the growing ones, preserving MoE -> KV -> mamba order within each phase.
Each pool still frees its own tensors first, so after the shrink phase the resident total is
sum(min(old, new)) and the grow phase rises monotonically to sum(new): the transient
peak never exceeds max(current, target), both of which already fit. All-grow and
all-shrink requests keep today's behavior.

Tests

New tests/engine/test_rebuild_runtime_cache_ordering.py: the issue's scenario (fake pools
under a shared cap, moe 40/kv 50 -> 60/30, cap 100) raises torch.OutOfMemoryError from the
MoE-first resize on unfixed main after the fit check passes, and completes after the fix with
peak == target total and the KV realloc ordered before the MoE realloc; plus reverse-direction
and rollback-restore guards. test_every_kv_pool_answers_the_sizing_surface gains the hook.

tests/engine/ + tests/kvcache/test_kv_cache_rebuild.py + tests/kvcache/test_pool_sizing_surface.py:
136 passed, 2 pre-existing env-gated skips. DSV4 / hybrid-SWA / V4.1 pool suites: 47 passed.

Hardware

RTX 5090, nvidia/Qwen3.8-Flash-Next-NVFP4 (offload MoE, moe 3400 slots / kv 317440 tok /
1.97 GiB free after CUDA graphs), command ft ctl cache rebuild --moe 4300 --kv 160k:

The mid-rebuild-OOM rollback failure itself remains #526.

KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 10, 2026
… resizes before growing ones in live rebuild
KarrAcaRn pushed a commit to KarrAcaRn/FreeToken-ByAI that referenced this pull request Oct 10, 2026
…ake engine next's dflash and graph-runner fields

Assisted-by: Claude Opus 5.5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Live cache rebuild grows MoE before shrinking KV/mamba, so a resize that passes the fit check can OOM

1 participant