chore(execution): parse memory sizes with one shared grammar - #1573
Merged
Merged
Conversation
`MLXCEL_MEMORY_LIMIT` was read by two parsers with two grammars. `parse_memory_size` in `src/execution/runtime.rs` served the allocator cap and accepted `4G`, `4GB`, `512M`, `512MB` and plain bytes; the memory-estimation preflight in `src/execution/memory_estimate.rs` carried its own `parse_optional_memory_size_bytes` plus `parse_scaled_memory_size`, which took `GB` and `MB` only. So `MLXCEL_MEMORY_LIMIT=4G` capped the allocator and was silently dropped by `mlxcel inspect` and `--estimate-memory`, which then reported availability from the machine's total memory instead. `inspect` runs before runtime bring-up, so the MLX-limit fallback is zero there too and nothing caught the divergence. `parse_memory_size` is now the one grammar and is `pub(crate)`, so the preflight resolves a string to the same number the allocator cap will apply. It returns `u64`, adds `K`/`KB` alongside the existing `M`/`MB` and `G`/`GB`, rejects a negative, `NaN` or infinite numeric part instead of casting it to zero, and states the overflow saturation explicitly rather than leaning on the implicit saturating float-to-int cast that both old parsers relied on. Every spelling that parsed before still parses to the same value; the three resolvers keep their own `0`/`none`/empty (and `max`) handling and narrow to `usize` at the MLX setter boundary through `clamp_to_usize`. One outcome does change: a garbage suffixed value such as `MLXCEL_WIRED_LIMIT=-1GB` used to cast to zero and silently disable the wired limit, and now falls back to `gpu_max_memory_size()` the way `MLXCEL_WIRED_LIMIT=abc` always did. The preflight's `parse_optional_memory_size_bytes` keeps only its unset check and delegates the rest, and `parse_scaled_memory_size` is gone. The blast radius is the reason the old spellings are pinned first: `parse_memory_size` also governs `MLXCEL_WIRED_LIMIT` and `MLXCEL_CACHE_LIMIT`, so a grammar or return-type change moves allocator caps for every run. `parse_memory_size_gb`, `_mb`, `_bytes`, `_fractional_gb` and `_invalid` stay as they were, and `parse_memory_size_accepts_every_suffix_spelling`, `_fractional_is_exact_floor`, `_rejects_garbage` and `_saturates_instead_of_wrapping` are added beside them. On the preflight side, `available_memory_honors_short_suffix_env_limit` drives `MLXCEL_MEMORY_LIMIT=512M` through `estimate_total_memory` and asserts the same 512 MiB the `512MB` sibling asserts, and `parse_optional_memory_size_accepts_the_runtime_grammar` pins `4G` == `4GB` == `4gb`. Documentation follows the code: the three size-valued rows in `docs/environment-variables.md` list every accepted suffix, a new paragraph under the table states the shared grammar once (binary units, fractional allowed on a suffixed value, floored, integer-only for a bare byte count), and the `mlxcel --help` environment block now reads `GB/G, MB/M, KB/K, or bytes`. Validated with `cargo test --profile test-fast --features metal,accelerate --lib execution::` (126 passed, 0 failed), `cargo clippy --profile test-fast --lib --tests --features metal,accelerate -- -D warnings`, and `cargo fmt --all -- --check`. Refs #1317
inureyes
force-pushed
the
refactor/issue-1317-one-memory-size-grammar
branch
from
September 2, 2026 00:35
3a396d0 to
078cf8a
Compare
3 of 5 tasks
Merged
7 of 17 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
MLXCEL_MEMORY_LIMITwas read by two parsers with two grammars, soMLXCEL_MEMORY_LIMIT=4Gcapped the MLX allocator and was silently ignored by the memory-estimation preflight, which then reported availability from the machine's total memory. This keeps one parser insrc/execution/runtime.rs, exposes it tosrc/execution/memory_estimate.rs, and pins the grammar with tests on both sides.What changed
src/execution/runtime.rs:parse_memory_sizebecomespub(crate) fn parse_memory_size(&str) -> Option<u64>and is the single grammar forMLXCEL_MEMORY_LIMIT,MLXCEL_WIRED_LIMITandMLXCEL_CACHE_LIMIT. It addsK/KBbeside the existingM/MBandG/GB, rejects a negative,NaNor infinite numeric part instead of casting it to zero, floors a fractional value, and states the overflow saturation explicitly rather than leaning on the implicit saturating float-to-int cast that both old parsers relied on. Every spelling that parsed before parses to the same value.src/execution/runtime.rs: the three resolvers keep their own0/none/ empty (andmax) handling and narrow tousizeat the MLX setter boundary through a newclamp_to_usize, soRuntimeSetupand themlxcel_coresetters see the types they saw before.src/execution/memory_estimate.rs:parse_scaled_memory_sizeis deleted andparse_optional_memory_size_byteskeeps only its unset check before delegating tocrate::execution::runtime::parse_memory_size(...).filter(|b| *b > 0). The preflight and the allocator cap now resolve one string to one number.src/execution/runtime_tests.rs: the five existingparse_memory_size_*tests stay as regression pins for the old spellings, andparse_memory_size_accepts_every_suffix_spelling,parse_memory_size_fractional_is_exact_floor,parse_memory_size_rejects_garbageandparse_memory_size_saturates_instead_of_wrappingare added beside them.src/execution/memory_estimate.rstests:available_memory_honors_short_suffix_env_limitdrivesMLXCEL_MEMORY_LIMIT=512Mthroughestimate_total_memoryand asserts the same 512 MiB its512MBsibling asserts, andparse_optional_memory_size_accepts_the_runtime_grammarpins4G==4GB==4gb.docs/environment-variables.md: the three size-valued rows list every accepted suffix, and a paragraph under the table states the shared grammar once (binary units, fraction allowed on a suffixed value, floored, integer-only for a bare byte count).src/main.rs: the--helpenvironment block now readsGB/G, MB/M, KB/K, or bytesforMLXCEL_WIRED_LIMITandMLXCEL_MEMORY_LIMIT.Blast radius
parse_memory_sizealso governsMLXCEL_WIRED_LIMITandMLXCEL_CACHE_LIMIT, so the grammar and return type here move allocator caps for every run. The old spellings are pinned by the pre-existing tests, which are unchanged apart from theusizetou64literal inparse_memory_size_fractional_gb. One outcome changes, and only for garbage that used to cast to zero: a negative or non-finite suffixed value (-1GB,NaNGB,infGB) now returnsNonerather thanSome(0). ForMLXCEL_MEMORY_LIMITandMLXCEL_CACHE_LIMITthat is the same end state (both mapped a zero to unset already); forMLXCEL_WIRED_LIMITit means-1GBfalls back togpu_max_memory_size()instead of silently disabling the wired limit, which is whatMLXCEL_WIRED_LIMIT=abcalways did. Neither parser ever wrapped: Rust's float-to-intascast has saturated since 1.45, so1e30GBresolved to the maximum before this change too.resolve_paged_slab_blocksinmemory_estimate.rsis untouched (issue #1137 owns it).Test plan
cargo test --profile test-fast --features metal,accelerate --lib execution::passes: 126 passed, 0 failed, including all 13 parser and preflight tests named above.cargo clippy --profile test-fast --lib --tests --features metal,accelerate -- -D warningspasses.cargo fmt --all -- --checkpasses.python3 scripts/ci/check_cross_repo_refs.pypasses.Acceptance commands for the release binary
Note that
mlxcel inspecttakes the model through-m/--model, not positionally, and the checkpoint present locally ismodels/mlx/qwen3-0.6b-4bit; the issue body'smlxcel inspect models/qwen3-0.6b-4bitis wrong on both counts.Expected: the first three print an identical
Available:line reading 4.00 GB, and the fourth prints the machine figure (the host's unified memory, so far larger). Onmainthe4Gand4096Minvocations print the machine figure instead, which is the defect.The allocator side, to confirm both readers agree on one string:
Expected: a
MLX allocator memory limit: 4.0 GB (MLXCEL_MEMORY_LIMIT)line at startup and normal generation. Onmainthat line already reads 4.0 GB for4G, which is exactly the half that the preflight disagreed with.Closes #1317