Repository navigation
fix(scheduler): stop the swa prefill commit from freeing pages the request still reads - #651
Open
ischencheng wants to merge 1 commit into
Open
ischencheng wants to merge 1 commit into
ischencheng wants to merge 1 commit into
Conversation
…quest still reads
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 10, 2026
…ll commit from freeing pages the request still reads
KarrAcaRn
pushed a commit
to KarrAcaRn/FreeToken-ByAI
that referenced
this pull request
Oct 10, 2026
…d tombstone through swa pool pressure next keeps the finished prompt head's swa (FlashML-org#488), so the finish no longer tombstones it; ensure_swa_slots does, and the test fails without the fix. Assisted-by: Claude Opus 5.5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #204.
_cache_req_swacan free pages that the committing request is still reading. It happens when the request's freshly prefilled span runs through a tombstone that another running request still full-locks:trim_head_swatombstones the head (orevict_swadoes it under pool pressure).match_prefixstops at the tombstone, so B prefills S itself into its own pages.insert. For a tombstone withref_count > 0,insertkeeps the tree's slots and returns B's copy of S infreed, which is then freed from both pools. The re-point after it only coversm.cached_len, and the re-match truncates to 0 again because the live run after the tombstone (B's tail) is shorter than the window. So B's row still names pages on the free list, and its in-window positions map to swa slot 0.free_slotsthat Concurrent tool-calling load kills the worker: duplicate pages in the full KV free list (root cause), surfacing as SWA-slot leak/double-free #204 measured, and the idlecheck_integrityfails withSWA-slot leak/double-free.The radix and hybrid paths re-point the deduped span at the tree's pages. That doesn't work here: the tree's copy has no swa, and B's window still covers the end of S. So this skips the unfinished commit when
insertwould take that branch, like the hybrid path does for a chunk with no tracked boundary. B keeps its own pages and its admission handle, and the finish commit inserts and frees the duplicates once. The finish path is unchanged.SWARadixCache.has_locked_tombstoneis a read-only walk with the same condition as that branch ofinsert.The cost is that in this case B holds its own copy of S until it finishes instead of sharing the tree's. Those are pages it already allocated for its prefill.
test_swa_unfinished_commit_keeps_its_pages_under_a_locked_tombstonebuilds the C/A/B sequence on a realCacheManagerandSWARadixCache(window 16), next to the radix and hybrid versions of the same test. On main, 64 positions of B's row are on the free list and 8 of its 16 in-window positions map to swa slot 0. After both requests finish,free_slotsholds 64 duplicates andcheck_integrityraisesSWA-slot leak/double-free: free(191) + tree(72) != capacity(255).Tested on an NVIDIA L4 (sm_89), driver 580.95.05, CUDA 13.0, torch 2.11.0+cu130, on top of 0781324:
pytest tests/scheduler/test_commit_repoints_page_table.py(this PR's file)pytest tests/scheduler -m "not slow"pytest tests/kvcache -m "not slow"End to end on an L40S (sm_89), same driver and torch, with
openai/gpt-oss-20b(sliding window 128):ft serve --model <snapshot> --max-seq-len-override 8192, which resolves to triton attention, offload MoE,swa_radixand page size 1. The script below sends, at temperature 0: C = S + 261 tokens decoding 1500 tokens, A = S + 514 tokens withmax_tokens=1once C is decoding, then B = S + 14 tokens. S is 2558 tokens. The scheduler log shows A reusing all 2558 tokens of S and B reusing none.--cache-type naiveWe/Thelines (differs run to run)check_integrity:SWA-slot leak/double-free: free(192409) + tree(781) != capacity(193175), thenBackend worker is gone and cannot be restarted; stopping the API servermain was run in 2 containers and this PR in 3, with the same outcome each time (the free/tree numbers above are from the first main run). B served alone on a fresh server also answers G-84, but its text drifts from the runs above after a few tokens. The naive run drifts from it in the same way, so that comes from B being served next to C, not from the prefix cache.
repro
#488 would remove the trim trigger, but not the
evict_swaone. Until this lands,--cache-type naiveavoids it, as the last column shows.