Skip to content

feat(vlm): add MiniMax-M3-VL multimodal support #764

Description

@inureyes

Summary

Add the vision-language variant of MiniMax-M3 (model_type: "minimax_m3_vl"). Depends on the MiniMax-M3 text port (#763).

Architecture specification

  • Vision tower: CLIP-style ViT (hidden 1280, 32 layers, LayerNorm plus GELU, a pre-layernorm whose checkpoint key is spelled pre_layrnorm), patch_size 14, with native-resolution packing: image_grid_thw, cu_seqlens variable-length attention, 3D RoPE over (t, h, w) axes, spatial_merge_size 2, temporal_patch_size 2. mlxcel already implements this packing pattern in src/vision/qwen2_5_vl.rs and src/vision/qwen3_vl.rs; reuse the cu_seqlens and RoPE helpers.
  • Projection: two stages, a per-patch multi_modal_projector (projection_dim 6144) followed by a patch_merge_mlp that folds spatial_merge_size^2 patches into the text hidden size. Visual tokens merge into the sequence by reshaping over the merge grid.
  • Processor: dynamic resolution (smart resize to factor-aligned dims, patchify into (grid_t, grid_h/m, m, grid_w/m, m)), emitting pixel_values plus image_grid_thw; video is analogous. Config block img_token_compression_config.{spatial_merge_size, temporal_patch_size}.
  • Token ids: image 200025, video 200026, vision start 200029, vision end 200030. Placeholder count per image is grid_t * (grid_h/2) * (grid_w/2).

Implementation plan

  1. src/vision/minimax_m3_vl.rs (tower plus projectors) and a processor in src/vision/processors/; reuse the qwen2_vl helper functions where they match.
  2. Loader arm in src/loading/ handling the checkpoint's actual dtype/layout split between quantized text weights and plain vision weights (verify against the real checkpoint, following the pattern of prior VLM ports).
  3. Registration in vision detection, model_metadata, generate_vlm summary, TP arch string, and docs/supported-models.md.
  4. Tests: grid math unit tests, placeholder count test, and a real-image smoke test.

Acceptance criteria

  • A real MiniMax-M3-VL checkpoint answers image questions correctly through CLI and server.
  • Placeholder expansion count matches the processor grid math for non-square images.
  • Lints and format are clean.

Validation

Real checkpoint, at least two images with distinct contents (object identification) plus one multi-image prompt. If no public VL checkpoint is available yet, keep this issue blocked and note the availability check in a comment.

Effort: high. Blocked by #763.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions