Skip to content

feat: fold sift bench into the binary and read safetensors - #3

Merged
maxgfr merged 2 commits into
mainfrom
phase45-bench-safetensors
Aug 2, 2026
Merged

maxgfr merged 2 commits into
mainfrom
phase45-bench-safetensors

Conversation

@maxgfr

@maxgfr maxgfr commented Aug 2, 2026

Copy link
Copy Markdown
Owner

The last two open items. TODO.md is now 14 done, 3 open.

sift bench

The sift-bench crate is gone — probe and models had become sift engines and sift ls. The harness now drives Ollama as well as LM Studio.

Ollama is deliberately not a copy of the LM Studio driver. Its native API reports eval_count and eval_duration, so its figure excludes prompt processing and HTTP overhead; the OpenAI shape LM Studio speaks reports no decode duration at all, so that path is wall clock. Each run records which clock produced it rather than presenting two different measurements as one.

--engine auto refuses to choose when both servers are up: silently preferring one would label a measurement with the wrong engine.

Runs append to ~/.sift/bench.jsonl. Two hand-set efficiency factors decide every tok/s figure sift prints, one still a guess — a measurement kept in a terminal scrollback cannot replace either.

Safetensors

fit, route and inspect read it. Each tensor's size is the difference of its data_offsets, never dtype × element count — same reasoning as reading real headers instead of shipping a catalog: an FP8 or MXFP4 variant this code has never heard of is sized exactly right by arithmetic that cannot go stale.

The header carries no architecture. GGUF keeps block_count and head_count_kv in its own header; a safetensors repo keeps them in config.json under different names. Rather than branch inside the KV arithmetic — where a mistake is most expensive — both formats normalise into ArchFacts, and ModelShape owns lightweight entries instead of borrowing GGUF tensors.

Three things that would have been quietly wrong:

  • route dispatches on format. mlx-lm reads safetensors, llama.cpp does not. Assuming GGUF recommended an engine that cannot open the file.
  • The download glob carries the extension. A .gguf pattern on a safetensors set matches nothing, and the user finds out only after the command runs.
  • MoE experts are counted by scanning names, not probing indices. A shard can hold experts 40–63 of a layer and nothing else; probing from zero stops at the first gap and reports a fraction of the model.

Verified on allenai/OLMoE-1B-7B-0924-Instruct:

  quant               size     fits    bpw    GB/token  est tok/s
  model           13.84 GB    tight  16.00       2.564         13  moe, 3 parts
  download   huggingface-cli download … --include "model-*.safetensors"

16.00 bpw is BF16 exactly — the arithmetic checks out against a known quantity.

130 tests, clippy clean on macOS, Linux and Windows targets.

maxgfr added 2 commits August 2, 2026 17:31
The last two open items in the plan.

**`sift bench`.** The `sift-bench` crate is gone: `probe` and `models` had
become `sift engines` and `sift ls`, and the only thing left was the harness.
It now drives Ollama as well as LM Studio.

Ollama is not a copy of the LM Studio driver on purpose. Its native API reports
`eval_count` and `eval_duration` — tokens decoded and nanoseconds spent
decoding, as the runtime counted them — which excludes prompt processing, HTTP
overhead and curl's startup. The OpenAI shape LM Studio speaks reports no
decode duration at all, so that path is wall clock. Each run records which
clock produced it rather than presenting two different measurements as one.

`--engine auto` refuses to choose when both servers are up. Silently preferring
one would label a measurement with the wrong engine, and that is a number
someone goes on to publish.

Runs are appended to `~/.sift/bench.jsonl`. Two hand-set efficiency factors
decide every tok/s figure sift prints, one still a guess; a measurement kept
only in a terminal scrollback cannot replace either.

**Safetensors.** `fit`, `route` and `inspect` now read it. Eight bytes of
length, then JSON, then one entry per tensor — and each tensor's size is the
difference of its `data_offsets`, never dtype times element count. That is the
same reasoning as reading real headers instead of shipping a catalog: an FP8 or
MXFP4 variant this code has never heard of is sized exactly right by
arithmetic that cannot go stale.

What the header does not carry is architecture. GGUF keeps `block_count` and
`head_count_kv` in its own header; a safetensors repo keeps them in
`config.json` under different names. Rather than branch inside the KV
arithmetic — the place where a mistake is most expensive — both formats now
normalise into `ArchFacts`, and `ModelShape` owns lightweight entries instead
of borrowing GGUF tensors.

`route` dispatches on format. mlx-lm reads safetensors and llama.cpp does not,
so assuming GGUF recommended an engine that cannot open the file. The download
glob carries the extension through for the same reason: a `.gguf` pattern on a
safetensors set matches nothing, and the user finds that out only after the
command runs.

MoE detection handles both layouts. GGUF stacks a layer's experts into one
tensor; safetensors keeps one per expert, and they are counted by scanning
names rather than probing indices — a shard can hold experts 40 to 63 of a
layer and nothing else, and probing from zero would stop at the first gap.

Verified on allenai/OLMoE-1B-7B-0924-Instruct: 13.84 GB, 16.00 bpw — BF16
exactly — MoE detected across 3 parts, KV sized from config.json.
`route` and `fit` were switched over; the other three still assumed GGUF, so
`sift inspect` on a safetensors file failed with "not a GGUF file: magic was
0x000228e0" — which is the length prefix of a perfectly valid header.

All five now go through `Source::read_shape`. GGUF-only details appear only
where they exist: the format version, and the layer-0 expert range, which needs
byte offsets that a merged or non-GGUF shape cannot provide.

`ls` also scanned for `.gguf` alone. MLX models ship as safetensors and live in
the same libraries, so on any machine with one the "every local model" table
was quietly showing a subset.
@maxgfr
maxgfr merged commit a6ec24f into main Aug 2, 2026
4 checks passed
@maxgfr
maxgfr deleted the phase45-bench-safetensors branch August 2, 2026 15:35
github-actions Bot pushed a commit that referenced this pull request Aug 2, 2026
# [0.3.0](v0.2.0...v0.3.0) (2026-08-02)

### Features

* fold sift bench into the binary and read safetensors ([#3](#3)) ([a6ec24f](a6ec24f))
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown

🎉 This PR is included in version 0.3.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant