Skip to content

Layered prefill chunk size - #240

Merged
stikves merged 10 commits into
apple:mainfrom
stikves:sukru/layered-prefill-chunk-size
Sep 11, 2026
Merged

stikves merged 10 commits into
apple:mainfrom
stikves:sukru/layered-prefill-chunk-size

Conversation

@stikves

@stikves stikves commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Replace the single COREAI_CHUNK_THRESHOLD environment variable with a layered override system for prefill chunking, exposing two independent settings:

  • prefillChunkSize — tokens processed per prefill chunk.
  • prefillChunkThreshold — minimum prompt length (in tokens) that triggers chunking; prompts at or below it are processed in a single pass.

Resolution order

Each setting is resolved by the first source that supplies a value, highest precedence first:

  1. CLI flag — --chunk-size / --chunk-threshold
  2. Foundation Models API — CoreAILanguageModel(prefillChunkSize:prefillChunkThreshold:)
  3. metadata.json — prefill_chunk_size / prefill_chunk_threshold
  4. Deprecated COREAI_CHUNK_THRESHOLD environment variable (emits a one-time warning)
  5. Memory-based default

This precedence applies uniformly across the CLI tools (llm-runner, llm-server, benchmark), the Foundation Models API, and both the text and vision-language paths.

Memory-based default

The default chunk size scales with installed RAM:

Installed RAM Default chunk size
≤ 24 GB 2048 tokens
> 24 GB 4096 tokens

The chunk threshold defaults to 2× the resolved chunk size.

This default applies to the dynamic (GPU) engine. The static-shape (Neural Engine) path derives its chunk width from the compiled model's fixed query-length buckets and does not consult this value.

Deprecation

COREAI_CHUNK_THRESHOLD continues to work and emits a one-time deprecation warning. Its value is applied as the chunk size (with the threshold set to the same value). Prefer --chunk-size or the prefill_chunk_size metadata key.

Test plan

  • ChunkingConfigTests covers layered resolution (override > env var > default), metadata.json decoding, invalid/zero overrides, Codable round-trip (overrides are runtime-only), the deprecated env var, and propagation through the VLM config.
  • CoreAILanguageModels, llm-runner, llm-server, and benchmark build clean.

@stikves stikves changed the title Sukru/layered prefill chunk size Layered prefill chunk size Sep 7, 2026
@stikves
stikves force-pushed the sukru/layered-prefill-chunk-size branch from 92f68d5 to cfe1919 Compare September 7, 2026 04:52
Replace the single COREAI_CHUNK_THRESHOLD env var with a proper layered
override system for prefill chunking. Two knobs: prefillChunkSize (tokens
per chunk) and prefillChunkThreshold (minimum prompt tokens to trigger
chunking).

Resolution order: CLI flag > FM API init > metadata.json > deprecated
env var > memory-based default.

Memory-based defaults scale chunk size with available RAM:
  <24 GB → 2048, ≤64 GB → 4096, ≤128 GB → 8192, >128 GB → 16384

The deprecated COREAI_CHUNK_THRESHOLD env var still works with a
one-time warning.
@stikves
stikves force-pushed the sukru/layered-prefill-chunk-size branch from 49b7dc6 to 93fc6d3 Compare September 9, 2026 06:21
stikves and others added 3 commits September 8, 2026 23:24
The layered chunk resolution only reached the engine on the CoreAIRunner/FM
path. The llm-runner, llm-server, and benchmark tools built EngineOptions
straight from CLI flags, so metadata.json prefill_chunk_size / _threshold were
ignored. Resolve them as CLI-flag ?? metadata before constructing EngineOptions,
keeping the value optional so the deprecated env var and memory-based default
still apply when neither is set.

Separately, CoreAISequentialVLMEngine never applied EngineOptions chunk
overrides to its base config, so --chunk-size had no effect for vision models.
Apply the overrides at engine init, mirroring EngineFactory, which covers both
the CLI and FM VLM construction sites; wire metadata into the FM VLM path too.
@stikves
stikves requested review from carinapeng, kevchengcodes and tjia1818 and removed request for carinapeng and tjia1818 September 10, 2026 01:25
@stikves stikves self-assigned this Sep 10, 2026
@stikves
stikves marked this pull request as ready for review September 10, 2026 01:26
Comment thread swift/Sources/Tools/llm-server/LLMServerMain.swift Outdated
stikves and others added 4 commits September 11, 2026 10:41
- ModelConfig: emit deprecation warning in the prefillChunkThreshold env-var
  branch so COREAI_CHUNK_THRESHOLD usage warns consistently with prefillChunkSize
- CLI help: clarify --chunk-size default text (128 is suggested for MoE models)
  across llm-runner, llm-server, and benchmark
@stikves
stikves merged commit 0d6c0bf into apple:main Sep 11, 2026
3 checks passed
@stikves
stikves deleted the sukru/layered-prefill-chunk-size branch September 11, 2026 23:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants