Skip to content

feat(kvcache): add ISO3/ISO4 KV-cache quantization (--kv-cache-iso) - #398

Open
AsmanovLev wants to merge 3 commits into
FlashML-org:mainfrom
AsmanovLev:iso-kv-cache
Open

AsmanovLev wants to merge 3 commits into
FlashML-org:mainfrom
AsmanovLev:iso-kv-cache

Conversation

@AsmanovLev

@AsmanovLev AsmanovLev commented Sep 5, 2026 •

Copy link
Copy Markdown

What

Optional IsoQuant quantization of the paged KV cache for plain
full-attention models, ported bit-exact from the llama.cpp fork
llama-cpp-turbo-planar-iso
(ggml ISO3_0/ISO4_0):

  • --kv-cache-iso iso3: 3.125 bits/value (50 B per 128-value block) - 5.1x smaller KV than bf16
  • --kv-cache-iso iso4: 4.25 bits/value (68 B per block) - 3.8x smaller

Per 128-value block of a head vector: L2-normalize, rotate each 4D group by a
fixed unit quaternion, quantize to the nearest Lloyd-Max centroid, store the
corrected norm ||x||/||centroids|| (norm-preserving dequant).

The flag forces --attention-backend iso: a new ISOKVCache pool (packed
uint8 slabs, drop-in MHAKVCache replacement incl. hybrid GDN layer remap) plus
custom CUDA kernels. Decode is quantize-on-write; extend attends the packed
prefix + bf16 new tokens and packs them after attention (deferred: prefill
never consumes its own quantized KV). Extend runs the packed prefix through a
one-time dequant into a bf16 scratch served by the existing tiled triton
kernel (custom CUDA extend kernel remains as fallback for empty prefix /
oversized scratch, FREETOKEN_ISO_SCRATCH_MB, default 128 MiB).

Why

On small-VRAM GPUs the KV cache - not the weights - caps the usable context.

Hardware:

  • RTX 3060 Laptop 6 GB (sm_86), driver 610.57.04, CUDA 13.3
  • I5-10500H, 32 GB DDR4

Checkpoint:

  ft serve --model-path <ckpt> --kv-cache-iso iso3 --memory-ratio 1.0 \
    --max-running-req 1 --num-tokens 81920 --moe-cache-size 256 \
    --disable-moe-prefill-overlap --cuda-graph-max-bs 0 --cache-type naive
main (bf16 KV) branch (iso3)
KV bytes/token (10 full-attn layers) 20,480 4,000
This model on 6 GB does not start (negative KV budget) 65k-80k token context
Decode - ~7-9 tok/s
Prefill - ~213 tok/s, flat with prefix length

Prefill A/B within the branch (custom CUDA extend kernel vs dequant+triton,
same 30k-token prompt): 3m42s -> 2m21s; throughput no longer degrades with
prefix length (~146 -> ~213 tok/s at 30k).

Correctness

  • quantize/dequantize bit-exact vs the fork's unmodified C reference
    (golden vectors: tests/kernels/iso_golden.npz); CUDA pack/unpack kernels
    byte-identical to the torch reference
  • attention kernels vs dequantized-oracle attention: cosine > 0.9999
  • 31 new tests (tests/kernels/test_iso_{quant,store_kernel,attention}.py,
    tests/kvcache/test_iso_pool.py); full tests/kvcache + tests/engine +
    tests/kernels suite passes (597 passed, 21 skipped)

Caveats

  • Quality is model-dependent: models with strong outlier K channels (dense
    Qwen3) degrade noticeably at iso3 (the 4D-group rotation cannot gaussianize
    channel spikes). Documented in docs/cli.md. Validated well-behaved on the
    Qwen3.5-MoE hybrid above (coherent 30k-context answers).
  • Plain full-attention pools only (no SWA/MLA/DSA/BSA/QSA/DSV4), bf16 KV
    input, head_dim % 128 == 0 (validated at config time with clear errors).

IsoQuant paged KV storage for plain full-attention models: per 128-value
block of a head vector, L2-normalize, rotate each 4D group by a fixed unit
quaternion, quantize to nearest Lloyd-Max centroid, store corrected norm.
Ported bit-exact from llama-cpp-turbo-planar-iso (its C reference is the
test oracle; golden vectors in tests/kernels/iso_golden.npz).

  iso3: 50 B / 128 values (3.125 bpw, ~5.1x smaller KV than bf16)
  iso4: 68 B / 128 values (4.25 bpw, ~3.8x smaller)

- kernel/csrc/jit/iso_{common.cuh,store.cu,attention.cu}: warp-cooperative
  pack/unpack + paged decode/extend attention kernels reading packed KV
  directly (extend = packed prefix + bf16 new tokens, deferred store).
- kvcache/iso_pool.py: ISOKVCache (uint8 slabs, packed rows, layer remap).
- attention/iso.py: iso backend (decode: quantize-on-write; extend:
  deferred store after attention so prefill never sees its own quantized
  K/V), CUDA-graph capture support.
- Engine wiring: --kv-cache-iso {off,iso3,iso4} forces --attention-backend
  iso; validates plain-FULL specs, bf16, head_dim % 128 == 0.
- kvcache/base.py: solve_num_pages assert now prints the budget numbers.

Validated on a 35B-A3B hybrid (Qwen3.5-MoE) at 64k-80k context on a 6 GB
RTX 3060. Quality note: models with strong outlier K channels (Qwen3 dense)
degrade at iso3 -- see docs/cli.md caveat.
…sh kernel

Prefill/extend on the ISO pool no longer runs the per-query-token custom
CUDA extend kernel (O(prefix) packed reads per query). Instead the packed
prefix is dequantized once per layer into a dense bf16 scratch and served by
the existing tiled triton extend kernel; the CUDA extend kernel remains as
the fallback for empty prefixes and oversized scratch (FREETOKEN_ISO_SCRATCH_MB,
default 128 MiB cap).

Measured on Qwen3.5-35B-A3B (Ornith NVFP4, iso3, RTX 3060 6GB): 30k-token
prefill 3m42s -> 2m21s; throughput no longer degrades with prefix length
(~213 tok/s at 30k vs ~146 before).
Copilot AI lite review requested due to automatic review settings September 5, 2026 20:24

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔵 Needs a closer look

It introduces new CUDA kernels and a new attention backend in a core execution path, which warrants careful human validation despite strong test coverage.

Pull request overview

This PR adds optional ISO3/ISO4 quantization for the paged KV cache, including a new ISO attention backend and CUDA pack/dequant/attention kernels, exposed via a new --kv-cache-iso server flag and documented in the CLI docs.

Changes:

  • Introduces ISOKVCache (packed uint8 KV pool) and ISO3/ISO4 reference quantization utilities.
  • Adds ISO decode/extend attention CUDA kernels and wires them into a new iso attention backend with config-time validation/forcing.
  • Adds targeted test coverage for quantization bit-exactness, store kernels, attention correctness, and pool sizing/accessors; updates CLI docs and server args.
File summaries
File Description
tests/kvcache/test_iso_pool.py New unit tests for ISOKVCache sizing, allocation, remapping, and store packing.
tests/kernels/test_iso_store_kernel.py New CUDA tests validating store/dequant kernels vs reference.
tests/kernels/test_iso_quant.py New tests for ISO reference quantize/dequant and golden-vector bit-exactness.
tests/kernels/test_iso_attention.py New CUDA tests validating ISO attention decode/extend vs dequantized oracle.
python/freetoken/server/args.py Adds --kv-cache-iso CLI flag with choices and help text.
python/freetoken/kvcache/iso_pool.py Adds ISOKVCache packed KV pool implementation.
python/freetoken/kvcache/base.py Improves KV-cache “not enough memory” assertion with detailed sizing context.
python/freetoken/kvcache/init.py Threads kv_cache_iso through KV pool creation and selects ISOKVCache when enabled.
python/freetoken/kernel/iso.py Adds ISO3/ISO4 reference quantization + JIT bindings for store/dequant and attention kernels.
python/freetoken/kernel/csrc/jit/iso_store.cu Adds CUDA pack/unpack (quantize-on-write) kernels for ISO KV.
python/freetoken/kernel/csrc/jit/iso_common.cuh Adds shared constants and device helpers for ISO quant/dequant.
python/freetoken/kernel/csrc/jit/iso_attention.cu Adds CUDA ISO attention decode/extend kernels (packed-prefix + bf16-extend).
python/freetoken/engine/engine.py Adds config-time validation for kv_cache_iso and forces/guards attention backend selection.
python/freetoken/engine/config.py Adds kv_cache_iso field to EngineConfig.
python/freetoken/attention/iso.py Adds IsoAttentionBackend integrating packed KV + CUDA kernels + dequant-to-triton fast path.
python/freetoken/attention/init.py Registers the new iso attention backend (FULL attention only).
docs/cli.md Documents --kv-cache-iso and adds iso to attention-backend list.
Review details
  • Files reviewed: 17/18 changed files
  • Comments generated: 1
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread python/freetoken/kvcache/iso_pool.py
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
@A-10-go-brrr

Copy link
Copy Markdown

Tried this branch on a 10 GB card. It works, and the context gain is real — but decode collapses with prefix depth, and several flags in your repro turn out to be 6 GB workarounds that cost a lot on larger cards. Data below.

Hardware

  • RTX 3080 10 GB (sm_86), driver 615.71.09, CUDA 13.4
  • AMD Ryzen 9 5900X 12-Core Processor (SMT off), 64 GB DDR4-3600 C16 (4×16 GB dual-rank, running 3200 CL16 / FCLK 1600)
  • PCIe 4.0 x16
  • CachyOS (Arch), Python 3.13
  • ft bench bw ceilings: CPU STREAM read 45.9 GB/s, PCIe H2D 25.8 / D2H 26.0 GB/s, nvfp4 CPU-MoE 42.5 vs PCIe-gather 26.3 (1.61x, backend offload)

Checkpoint

nvidia/Qwen3.6-35B-A3B-NVFP4 and ornith-ai/Ornith-1.5-35B-A3B-NVFP4 (Qwen3.5-MoE hybrid, 40 layers, 10 full-attn + 30 GDN), converted with ft checkpoint --moe-backend offload. Numbers below are Ornith 1.5.

Note: the Unsloth NVFP4 builds fail conversion with RuntimeError: Promotion for Float8 Types is not supported, attempted to promote Float8_e4m3fn and BFloat16 in loader.py:ct_bf16_fuse. The NVIDIA ModelOpt checkpoints convert cleanly. Unrelated to this PR, but worth knowing if others try to reproduce.

Your repro flags cost ~2x on a 10 GB card

Starting from your command and removing one flag at a time, short-prompt decode (~30 token prompt):

Config Decode (tok/s) TTFT VRAM
your flags as published 31 2484 ms 5.11 GB
drop --cuda-graph-max-bs 0 45–57 1697 ms 5.11 GB
also drop --disable-moe-prefill-overlap 59–63 834 ms 5.11 GB
also --cache-type radix, --moe-cache-size 2048 67–91 826 ms 7.90 GB

iso4 is 36% more bytes per block and 17% slower. If the cost were the unpack arithmetic (quaternion rotation, centroid lookup, norm correction) iso4 should be roughly flat or faster, since it's the same per-block work. Scaling with byte count instead suggests decode attention is bandwidth-bound re-reading the packed prefix — i.e. O(prefix) packed reads per generated token, the same pattern you fixed on the extend path in 0c5fe8f but on the decode path.

If that reading is right, the decode kernel might benefit from the same treatment: dequantize the prefix once into a bf16 scratch and let the existing paged decode kernel serve it, at least above some prefix-length threshold.

Where this leaves it

The context win is genuine — 212,992 tokens at 6.2 GB on a card that tops out near 24K with bf16 KV, which is exactly the problem the PR sets out to solve. For short-prompt interactive use with graphs and prefill overlap enabled it's also faster than the main branch on this hardware.

But at the context depths the branch exists to enable, decode is ~3 tok/s. That's workable for one-shot document summarization and not for anything iterative. Happy to run further tests if there's a specific variant you'd like measured.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants