Summary
Add the vision-language variant of MiniMax-M3 (model_type: "minimax_m3_vl"). Depends on the MiniMax-M3 text port (#763).
Architecture specification
- Vision tower: CLIP-style ViT (hidden 1280, 32 layers, LayerNorm plus GELU, a pre-layernorm whose checkpoint key is spelled
pre_layrnorm), patch_size 14, with native-resolution packing: image_grid_thw, cu_seqlens variable-length attention, 3D RoPE over (t, h, w) axes, spatial_merge_size 2, temporal_patch_size 2. mlxcel already implements this packing pattern in src/vision/qwen2_5_vl.rs and src/vision/qwen3_vl.rs; reuse the cu_seqlens and RoPE helpers.
- Projection: two stages, a per-patch
multi_modal_projector (projection_dim 6144) followed by a patch_merge_mlp that folds spatial_merge_size^2 patches into the text hidden size. Visual tokens merge into the sequence by reshaping over the merge grid.
- Processor: dynamic resolution (smart resize to factor-aligned dims, patchify into
(grid_t, grid_h/m, m, grid_w/m, m)), emitting pixel_values plus image_grid_thw; video is analogous. Config block img_token_compression_config.{spatial_merge_size, temporal_patch_size}.
- Token ids: image 200025, video 200026, vision start 200029, vision end 200030. Placeholder count per image is
grid_t * (grid_h/2) * (grid_w/2).
Implementation plan
src/vision/minimax_m3_vl.rs (tower plus projectors) and a processor in src/vision/processors/; reuse the qwen2_vl helper functions where they match.
- Loader arm in
src/loading/ handling the checkpoint's actual dtype/layout split between quantized text weights and plain vision weights (verify against the real checkpoint, following the pattern of prior VLM ports).
- Registration in vision detection, model_metadata, generate_vlm summary, TP arch string, and
docs/supported-models.md.
- Tests: grid math unit tests, placeholder count test, and a real-image smoke test.
Acceptance criteria
- A real MiniMax-M3-VL checkpoint answers image questions correctly through CLI and server.
- Placeholder expansion count matches the processor grid math for non-square images.
- Lints and format are clean.
Validation
Real checkpoint, at least two images with distinct contents (object identification) plus one multi-image prompt. If no public VL checkpoint is available yet, keep this issue blocked and note the availability check in a comment.
Effort: high. Blocked by #763.
Summary
Add the vision-language variant of MiniMax-M3 (
model_type: "minimax_m3_vl"). Depends on the MiniMax-M3 text port (#763).Architecture specification
pre_layrnorm), patch_size 14, with native-resolution packing:image_grid_thw, cu_seqlens variable-length attention, 3D RoPE over (t, h, w) axes,spatial_merge_size2,temporal_patch_size2. mlxcel already implements this packing pattern insrc/vision/qwen2_5_vl.rsandsrc/vision/qwen3_vl.rs; reuse the cu_seqlens and RoPE helpers.multi_modal_projector(projection_dim 6144) followed by apatch_merge_mlpthat foldsspatial_merge_size^2patches into the text hidden size. Visual tokens merge into the sequence by reshaping over the merge grid.(grid_t, grid_h/m, m, grid_w/m, m)), emittingpixel_valuesplusimage_grid_thw; video is analogous. Config blockimg_token_compression_config.{spatial_merge_size, temporal_patch_size}.grid_t * (grid_h/2) * (grid_w/2).Implementation plan
src/vision/minimax_m3_vl.rs(tower plus projectors) and a processor insrc/vision/processors/; reuse the qwen2_vl helper functions where they match.src/loading/handling the checkpoint's actual dtype/layout split between quantized text weights and plain vision weights (verify against the real checkpoint, following the pattern of prior VLM ports).docs/supported-models.md.Acceptance criteria
Validation
Real checkpoint, at least two images with distinct contents (object identification) plus one multi-image prompt. If no public VL checkpoint is available yet, keep this issue blocked and note the availability check in a comment.
Effort: high. Blocked by #763.