Layered prefill chunk size - #240
Merged
Merged
Conversation
stikves
force-pushed
the
sukru/layered-prefill-chunk-size
branch
from
September 7, 2026 04:52
92f68d5 to
cfe1919
Compare
Replace the single COREAI_CHUNK_THRESHOLD env var with a proper layered override system for prefill chunking. Two knobs: prefillChunkSize (tokens per chunk) and prefillChunkThreshold (minimum prompt tokens to trigger chunking). Resolution order: CLI flag > FM API init > metadata.json > deprecated env var > memory-based default. Memory-based defaults scale chunk size with available RAM: <24 GB → 2048, ≤64 GB → 4096, ≤128 GB → 8192, >128 GB → 16384 The deprecated COREAI_CHUNK_THRESHOLD env var still works with a one-time warning.
stikves
force-pushed
the
sukru/layered-prefill-chunk-size
branch
from
September 9, 2026 06:21
49b7dc6 to
93fc6d3
Compare
The layered chunk resolution only reached the engine on the CoreAIRunner/FM path. The llm-runner, llm-server, and benchmark tools built EngineOptions straight from CLI flags, so metadata.json prefill_chunk_size / _threshold were ignored. Resolve them as CLI-flag ?? metadata before constructing EngineOptions, keeping the value optional so the deprecated env var and memory-based default still apply when neither is set. Separately, CoreAISequentialVLMEngine never applied EngineOptions chunk overrides to its base config, so --chunk-size had no effect for vision models. Apply the overrides at engine init, mirroring EngineFactory, which covers both the CLI and FM VLM construction sites; wire metadata into the FM VLM path too.
stikves
requested review from
carinapeng,
kevchengcodes and
tjia1818
and removed request for
carinapeng and
tjia1818
September 10, 2026 01:25
stikves
marked this pull request as ready for review
September 10, 2026 01:26
tjia1818
reviewed
Sep 11, 2026
- ModelConfig: emit deprecation warning in the prefillChunkThreshold env-var branch so COREAI_CHUNK_THRESHOLD usage warns consistently with prefillChunkSize - CLI help: clarify --chunk-size default text (128 is suggested for MoE models) across llm-runner, llm-server, and benchmark
tjia1818
approved these changes
Sep 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Replace the single
COREAI_CHUNK_THRESHOLDenvironment variable with a layered override system for prefill chunking, exposing two independent settings:prefillChunkSize— tokens processed per prefill chunk.prefillChunkThreshold— minimum prompt length (in tokens) that triggers chunking; prompts at or below it are processed in a single pass.Resolution order
Each setting is resolved by the first source that supplies a value, highest precedence first:
--chunk-size/--chunk-thresholdCoreAILanguageModel(prefillChunkSize:prefillChunkThreshold:)metadata.json—prefill_chunk_size/prefill_chunk_thresholdCOREAI_CHUNK_THRESHOLDenvironment variable (emits a one-time warning)This precedence applies uniformly across the CLI tools (
llm-runner,llm-server,benchmark), the Foundation Models API, and both the text and vision-language paths.Memory-based default
The default chunk size scales with installed RAM:
The chunk threshold defaults to 2× the resolved chunk size.
This default applies to the dynamic (GPU) engine. The static-shape (Neural Engine) path derives its chunk width from the compiled model's fixed query-length buckets and does not consult this value.
Deprecation
COREAI_CHUNK_THRESHOLDcontinues to work and emits a one-time deprecation warning. Its value is applied as the chunk size (with the threshold set to the same value). Prefer--chunk-sizeor theprefill_chunk_sizemetadata key.Test plan
ChunkingConfigTestscovers layered resolution (override > env var > default),metadata.jsondecoding, invalid/zero overrides, Codable round-trip (overrides are runtime-only), the deprecated env var, and propagation through the VLM config.CoreAILanguageModels,llm-runner,llm-server, andbenchmarkbuild clean.