feat: fold sift bench into the binary and read safetensors - #3
Merged
Merged
Conversation
The last two open items in the plan. **`sift bench`.** The `sift-bench` crate is gone: `probe` and `models` had become `sift engines` and `sift ls`, and the only thing left was the harness. It now drives Ollama as well as LM Studio. Ollama is not a copy of the LM Studio driver on purpose. Its native API reports `eval_count` and `eval_duration` — tokens decoded and nanoseconds spent decoding, as the runtime counted them — which excludes prompt processing, HTTP overhead and curl's startup. The OpenAI shape LM Studio speaks reports no decode duration at all, so that path is wall clock. Each run records which clock produced it rather than presenting two different measurements as one. `--engine auto` refuses to choose when both servers are up. Silently preferring one would label a measurement with the wrong engine, and that is a number someone goes on to publish. Runs are appended to `~/.sift/bench.jsonl`. Two hand-set efficiency factors decide every tok/s figure sift prints, one still a guess; a measurement kept only in a terminal scrollback cannot replace either. **Safetensors.** `fit`, `route` and `inspect` now read it. Eight bytes of length, then JSON, then one entry per tensor — and each tensor's size is the difference of its `data_offsets`, never dtype times element count. That is the same reasoning as reading real headers instead of shipping a catalog: an FP8 or MXFP4 variant this code has never heard of is sized exactly right by arithmetic that cannot go stale. What the header does not carry is architecture. GGUF keeps `block_count` and `head_count_kv` in its own header; a safetensors repo keeps them in `config.json` under different names. Rather than branch inside the KV arithmetic — the place where a mistake is most expensive — both formats now normalise into `ArchFacts`, and `ModelShape` owns lightweight entries instead of borrowing GGUF tensors. `route` dispatches on format. mlx-lm reads safetensors and llama.cpp does not, so assuming GGUF recommended an engine that cannot open the file. The download glob carries the extension through for the same reason: a `.gguf` pattern on a safetensors set matches nothing, and the user finds that out only after the command runs. MoE detection handles both layouts. GGUF stacks a layer's experts into one tensor; safetensors keeps one per expert, and they are counted by scanning names rather than probing indices — a shard can hold experts 40 to 63 of a layer and nothing else, and probing from zero would stop at the first gap. Verified on allenai/OLMoE-1B-7B-0924-Instruct: 13.84 GB, 16.00 bpw — BF16 exactly — MoE detected across 3 parts, KV sized from config.json.
`route` and `fit` were switched over; the other three still assumed GGUF, so `sift inspect` on a safetensors file failed with "not a GGUF file: magic was 0x000228e0" — which is the length prefix of a perfectly valid header. All five now go through `Source::read_shape`. GGUF-only details appear only where they exist: the format version, and the layer-0 expert range, which needs byte offsets that a merged or non-GGUF shape cannot provide. `ls` also scanned for `.gguf` alone. MLX models ship as safetensors and live in the same libraries, so on any machine with one the "every local model" table was quietly showing a subset.
github-actions Bot
pushed a commit
that referenced
this pull request
Aug 2, 2026
# [0.3.0](v0.2.0...v0.3.0) (2026-08-02) ### Features * fold sift bench into the binary and read safetensors ([#3](#3)) ([a6ec24f](a6ec24f))
|
🎉 This PR is included in version 0.3.0 🎉 The release is available on GitHub release Your semantic-release bot 📦🚀 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The last two open items. TODO.md is now 14 done, 3 open.
sift benchThe
sift-benchcrate is gone —probeandmodelshad becomesift enginesandsift ls. The harness now drives Ollama as well as LM Studio.Ollama is deliberately not a copy of the LM Studio driver. Its native API reports
eval_countandeval_duration, so its figure excludes prompt processing and HTTP overhead; the OpenAI shape LM Studio speaks reports no decode duration at all, so that path is wall clock. Each run records which clock produced it rather than presenting two different measurements as one.--engine autorefuses to choose when both servers are up: silently preferring one would label a measurement with the wrong engine.Runs append to
~/.sift/bench.jsonl. Two hand-set efficiency factors decide every tok/s figure sift prints, one still a guess — a measurement kept in a terminal scrollback cannot replace either.Safetensors
fit,routeandinspectread it. Each tensor's size is the difference of itsdata_offsets, never dtype × element count — same reasoning as reading real headers instead of shipping a catalog: an FP8 or MXFP4 variant this code has never heard of is sized exactly right by arithmetic that cannot go stale.The header carries no architecture. GGUF keeps
block_countandhead_count_kvin its own header; a safetensors repo keeps them inconfig.jsonunder different names. Rather than branch inside the KV arithmetic — where a mistake is most expensive — both formats normalise intoArchFacts, andModelShapeowns lightweight entries instead of borrowing GGUF tensors.Three things that would have been quietly wrong:
routedispatches on format. mlx-lm reads safetensors, llama.cpp does not. Assuming GGUF recommended an engine that cannot open the file..ggufpattern on a safetensors set matches nothing, and the user finds out only after the command runs.Verified on
allenai/OLMoE-1B-7B-0924-Instruct:16.00 bpw is BF16 exactly — the arithmetic checks out against a known quantity.
130 tests, clippy clean on macOS, Linux and Windows targets.