Repository navigation
Decompose small kernel-depth 3D convs into 2D convs - #3785
Conversation
|
Gentle ping on this one — it's approved with all CI green. Happy to rebase onto the latest main if you'd like it current; otherwise it's ready to merge whenever you're batching. Thanks! |
|
@katlun-lgtm For changes that I'm not 100% sure I keep them opened for a while and ask other maintainers for a second view. |
|
Checking in on this one. It's been approved since July 3, CI went green on As I understand it the only thing holding it is the second opinion @zcbenz mentioned on July 7:
That's a reasonable ask for this change — it decomposes small kernel-depth 3D convs into 2D convs, so the decomposition boundaries are the part worth a second pair of eyes. It just hasn't come up since, and I'd rather it not quietly age out. If another maintainer can give that second view, I'm glad to answer questions, add cases to the test, or rerun the benchmarks on whatever shapes you'd want to see. |
3x3x3 (and other small kernel-depth) 3D convolutions run the generic 3D implicit-gemm kernel, which has no Winograd / 3x3-specialized path. As a result they are 2-5x slower than decomposing the same op into per-frame 2D convolutions (which do hit the tuned 2D dispatch). This shows up in video VAE decoders (CausalConv3d). Add small_kd_conv_3D_gpu: for small kernel depth, run KD 2D convs over zero-copy strided views of the input frames and weight depth-slices, and accumulate. Each 2D conv goes through dispatch_conv_2D_gpu, so 3x3 stride-1 taps get Winograd. A guard in dispatch_conv_3D_gpu takes this path only when it is valid and faster: input dilation 1, groups 1, N == 1, KD <= 7, depth stride and kernel-dilation 1, no depth padding, mod16 channels. Everything else falls through to the existing implicit gemm unchanged. Tests: the fast path is validated against the CPU reference (fp32, exact) across several shapes, plus fall-back cases (depth stride > 1, depth padding, non-mod16 channels). Benchmarked on an M3 Max (bf16): 41x120x208, C=512, 3x3x3 goes 1128ms -> 335ms (3.4x -> 1.0x vs the per-frame 2D decomposition); fp32-exact vs CPU.
237ac5f to
5d6d778
Compare
5d6d778 to
63145e1
Compare
* Return tuple in meshgrid (ml-explore#4229) * Add endpoint parameter to linspace (ml-explore#4184) Co-authored-by: Cheng <git@zcbenz.com> * Fix vmap of partition/argpartition dropping the kth argument (ml-explore#4116) * Fix nan_to_num replacing inf with 0 for float16 and bfloat16 (ml-explore#4222) Co-authored-by: codeAnqiang-ma <273298913+codeAnqiang-ma@users.noreply.github.com> Co-authored-by: Cheng <git@zcbenz.com> * Fix einsum not broadcasting batch dimensions in batched tensordot (ml-explore#4125) Co-authored-by: Cheng <git@zcbenz.com> * Dequantize in float32 (ml-explore#4241) * chore: Reject complex in erf and erfinv (ml-explore#4243) * Fix cpu compilation failure of abs with uint (ml-explore#4240) Co-authored-by: Cheng <git@zcbenz.com> * Fix quantize matrix multiplication floor issue (ml-explore#4251) * Only use MPI backend for world size > 1 (ml-explore#4210) * chore: Reject complex in expm1, sigmoid and arctan2 (ml-explore#4257) * Decompose small kernel-depth 3D convs into 2D convs (ml-explore#3785) Co-authored-by: katlun-lgtm <264247399+katlun-lgtm@users.noreply.github.com> Co-authored-by: Cheng <git@zcbenz.com> * Fix Metal sort of a view with a negative stride (ml-explore#4252) * Mirror the depth axis in the decomposed 3D conv when flipped (ml-explore#4277) * Fix Metal row reductions on negative-stride views (ml-explore#4267) Co-authored-by: Fu Xiaonan <214359569+FU-max-boop@users.noreply.github.com> * [CUDA] Fix custom kernel cache collision for same name, different source (ml-explore#4273) Co-authored-by: Cheng <git@zcbenz.com> * Fix ops rejecting integers larger than INT32_MAX (ml-explore#4255) Co-authored-by: Feli <feli@hnu.edu.cn> Co-authored-by: Cheng <git@zcbenz.com> * Fix var/std for complex numbers (ml-explore#4260) * Fix int32 overflow in conv padded input and pad shapes (ml-explore#4258) Co-authored-by: Cheng <git@zcbenz.com> * chore: Reject complex in remainder (ml-explore#4270) * chore: Compare the macOS SDK version as a version when gating JACCL (ml-explore#4286) * Clamp ring socket transfers so a payload of 2 GiB or more can be sent (ml-explore#4281) Co-authored-by: Cheng <git@zcbenz.com> * chore: Use normalize_axis_index in split/unstack/partition/topk (ml-explore#4288) * Remove grouped output in CI (ml-explore#4195) * [CUDA] Fix finding cuda 13 headers in JIT compilation (ml-explore#3995) * Refactor wheel building script (ml-explore#3818) * Make mx.compile cache erasing thread safe (ml-explore#4248) Co-authored-by: yentur <mr.yentur@gmail.com> * Add builds for free-threaded python (ml-explore#3812) * Fix int32 overflow in concatenate/repeat/kron (ml-explore#4303) * python: Widen list elements that do not fit in int32 to int64 (ml-explore#4305) * Propagate CPU errors to events (ml-explore#3742) Co-authored-by: Alessio Pollero <alessio.pollero@gmail.com> * Fix mx.arange dtype inference overflow regression (ml-explore#4324) * Add workflow to update pull request limit bypass list (ml-explore#4320) * Support head dimension 72 in Metal full attention (ml-explore#4330) * Patch bump to 0.32.2 (ml-explore#4333) * Preserve subnormal float values when casting to bool (ml-explore#4224) * python: Support assigning through a bare Ellipsis index (ml-explore#4314) * Fix divmod truncating the quotient for floats (ml-explore#4108) Co-authored-by: Cheng <git@zcbenz.com> * Add force_fused option to scaled_dot_product_attention (ml-explore#4185) * chore: Reject negative eps in the normalization layers (ml-explore#4312) * Bound GGUF metadata string/array values against the file mapping (ml-explore#4212) Co-authored-by: x14ngch3n <x14ngch3n@users.noreply.github.com> Co-authored-by: Cheng <git@zcbenz.com> * Read each K/V byte once in gqa-8 decode attention (ml-explore#4077) * Fix fft vmap and jvp for transforms over a subset of axes (ml-explore#4138) * Fix median dropping NaN (ml-explore#4146) * Fix the CPU scan over a size one axis with a padded stride (ml-explore#4139) Co-authored-by: Cheng <git@zcbenz.com> * chore: Validate the optimizer betas at construction (ml-explore#4310) Co-authored-by: Cheng <git@zcbenz.com> * `RMSNormVJP` backward writes a full `{n_rows, D}` `gw_temp` intermediate (ml-explore#4293) * [Bug]: add default none value to axis parameter of the take_along_axis (ml-explore#4357) Co-authored-by: Anastasiia Filippova <a_filippova@apple.com> * Add a fused full-attention path for head_dim 256 on NAX devices (ml-explore#3842) Co-authored-by: Cheng <git@zcbenz.com> * Update nanobind to 2.15.0 (ml-explore#4337) * Skip unnecessary simdgroup computations for quantised MOE matmuls on NAX (ml-explore#4352) * Add AI usage policy (ml-explore#4331) Co-authored-by: Jake Bowhay <60778417+j-bowhay@users.noreply.github.com> * Raise cpu stream errors from synchronize (ml-explore#4338) Co-authored-by: Cheng <git@zcbenz.com> * chore: Validate eps in Adam at construction (ml-explore#4361) Co-authored-by: Anastasiia Filippova <a_filippova@apple.com> * Bound winograd conv2d working set by tiling the batch (ml-explore#4102) Co-authored-by: Cheng <git@zcbenz.com> * Use a 32-row block in qmm_t_nax when one block covers all of M (ml-explore#4171) * chore: Deduplicate fftshift and ifftshift (ml-explore#4318) * Fix Log and Equal is_equivalent ignoring primitive state (ml-explore#4266) Co-authored-by: Cheng <git@zcbenz.com> * Stabilize reduced-precision InstanceNorm (ml-explore#4230) * chore: Normalize negative axes in sort and argsort (ml-explore#4332) * Clean up main thread compile cache before python interpreter shuts down (ml-explore#4373) * chore: Check malformed jaccl hostfile that miss rdma in pairs (ml-explore#4284) Co-authored-by: Cheng <git@zcbenz.com> * Round mxfp8 block scales up to avoid saturation (ml-explore#4353) Co-authored-by: Daniel Hiltgen <daniel.hiltgen@ollama.com> Co-authored-by: Cheng <git@zcbenz.com> * Add support for the __array_namespace_info__ (ml-explore#4334) * Stop a failed CUDA graph commit from poisoning the encoder (ml-explore#4356) Co-authored-by: Cheng <git@zcbenz.com> * Fix quantized kernels in JIT build (ml-explore#4372) Co-authored-by: Cheng <git@zcbenz.com> * Avoid zero work in stride-2 ConvTranspose3d (ml-explore#4343) * [CUDA] Ce fused kernel (ml-explore#3947) * Fix cpu exclusive scan for complex numbers (ml-explore#4272) Co-authored-by: Cheng <git@zcbenz.com> * Support Relocatable CUDA DLLs on Windows (ml-explore#4382) * Use cast_to for fused AsType in compiled Metal kernels (ml-explore#4351) Co-authored-by: katlun-lgtm <katlun@windyviews.com> Co-authored-by: Cheng <zcbenz@gmail.com> * python: Declare DLPackCompatible protocol members as methods (ml-explore#4384) * Fix quantizing sliced arrays (ml-explore#4381) * Fix einsum dropping a trailing empty subscript (ml-explore#4299) Co-authored-by: Cheng <git@zcbenz.com> * Add script to run python tests (ml-explore#4393) * Hold GIL in AttachedData destructor (ml-explore#4391) * Bound Metal buffer COUNT, not just bytes, in MetalAllocator The Metal allocator throws `[metal::malloc] Resource limit (N) exceeded` when num_resources_ (the live+cached Metal buffer COUNT) reaches resource_limit_ (the iogpu.rsrc_limit sysctl, default ~499000). Freed buffers are recycled into a size-keyed cache whose only trim is by BYTES (release_cached_buffers takes a bytes-to-free target, max_pool_size_ ~= physical RAM). Under churn with many distinct buffer shapes (varied prompt lengths, growing KV caches, multiple co-resident models) the cache fills with entries never reused at that exact size, so the COUNT climbs to the limit while byte usage stays modest and the byte trim never fires — the process crashes mid-inference on a machine with most of its RAM free. malloc() now also reclaims by count: when num_resources_ crosses a 90% high-water mark of resource_limit_, it clears the (pure-reuse) buffer cache so the count drops back to the live working set. Clearing the cache only costs re-allocation, never correctness, so the count limit becomes unreachable by any request mix or batching method while the existing byte limits keep total memory bounded. Adds get_num_resources()/get_resource_limit() to the public memory API (metal + no_gpu + cuda backends) so the count and its ceiling are observable from callers. Adds an MLX_RESOURCE_LIMIT env override that can only LOWER the ceiling (clamped to the OS limit, strictly validated) to exercise the trim deterministically and as an operator safety valve. * perf(mlx): opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder (#4) * perf(mlx): add opt-in Gemma 4 expert-QMM tile kernel with parallel descriptor builder Adds a distinctly-named expert QMM implementation for the Gemma 4 26B-A4B MoE production shapes, gated by MLX_GATHER_QMM_EXPERT_SLICES: - qmm_t_expert_impl: BM32 expert tile body (BM16 fallback rows) taking a private/by-value row count; the shared qmm_t_impl constant-address ABI and all ordinary gathered/batched/dense QMM routes are unchanged. - build_gemma4_sorted_expert_tiles_bm32: one 128-thread threadgroup replaces the reference design's single-GPU-thread serial builder; parallel expert-range binary search, Hillis-Steele scan, and strided upper-bound descriptor emission. - Selector runs after the NAX-first route and requires affine BF16 transposed inputs, 4-bit gs=64 weights, 128 experts, assignment counts of exactly 4096/8192/16384, and the exact gate/up or down rank-3 shapes; every miss keeps the legacy route. NAX engagement is non-engagement, never bypassed. - device.{h,cpp}: one-shot request resolution, nonthrowing dual-symbol AOT probe/prewarm, relaxed-atomic diagnostics (requested, aotAvailable, naxAvailable, hits, per-class fallbacks). - gpu_tests: exact-shape arithmetic parity, fallback, and counter invariant probes. Retention standing (2026-08-09 production matrix): opt-in experiment. Standalone profile dropped (prefill -10.2% vs bracket); paired weighted-unsort+R1 profile retained-final (prefill +1.8%, TTFT -7.5%, decode +3.3%, arrival E2E +12.0%). NOTE: this source post-dates the benchmarked binaries/metallib (post-measurement kernel-body edit); rebuild and re-verify before any performance claim. * fix(mlx): fail-safe sortedness check in gemma expert tile builder; counter/atomic hygiene Review-wave fixes for the R1 expert-QMM path: - N1 (sortedness trust): build_gemma4_sorted_expert_tiles_bm32 now verifies each thread's post-binary-search segment boundary against the generalized invariant indices[start - 1] < lid <= indices[start] (edge threads check their single neighbor), votes per simdgroup via simd_or, folds the votes through threadgroup memory, and on any violation retracts count[0] to 0 (tile kernel then early-returns) and records the violation in count[1]; the buffer ABI is unchanged (count index 1 was previously unused). try_gemma4_expert_qmm allocates the second count element, drains the encoder after the builder, and re-routes a retracted call to the order-agnostic legacy path instead of dispatching the tile kernel (zero count is unambiguous: the selector's assignment gate guarantees M is 4096/8192/16384). - N2 (route-condition duplication): the sorted-RHS gate literal that appeared (negated) in the diagnostics record and in the dispatch decision is now the shared static constexpr predicate takes_sorted_rhs_route, so future tuning of the 16/4 thresholds cannot desynchronize counter vs route. - N3 (per-call bias normalization): gather_qmm_rhs no longer spends ensure_row_contiguous on biases before classification reads the raw tensor's fields; normalization runs only inside the winning-route branch (hit semantics unchanged; the legacy block keeps its own normalization point and ordering). - N4 (armed_ data race): Gemma4ExpertQMMCounters::armed_ is now std::atomic<bool> with relaxed loads/stores in armed(), snapshot(), snapshot_and_disarm() (read-then-write order preserved) and clear_and_arm(); the class remains non-copyable, now enforced. * fix(mlx): make the R1 sortedness fail-safe sound; proper retract attribution F1: the per-expert boundary vote was a partial detector -- an inversion inside a segment used by no other expert's boundary could escape, so "re-route on any violation" overclaimed. build_gemma4_sorted_expert_tiles_bm32 now also runs a strided adjacent-pair scan: thread lid checks indices[i-1] <= indices[i] for i = lid+1; i < M; i += 128, covering every adjacent pair in [1, M) exactly once (1..128 iterations at the reachable M in {4096,8192,16384}). Adjacent-pair monotonicity is transitive, so a clean scan is a sound and complete sortedness oracle; it folds into the same simd_or/threadgroup vote and the same retract (count[0]=0, count[1]=1). The boundary checks stay as cheap, precise diagnostics. F2: retracts were write-only in count[1] and surfaced as fallback_metallib_unavailable -- misattribution in the only observable surface. A dedicated fallback_sortedness_retracted counter now rides the GemmA4 route counters and the C diagnostics ABI (sizeof 80 -> 88, new uint64 at offset 80; existing offsets unchanged). try_gemma4_expert_qmm returns the route class: count[0]==0 with count[1]==1 records fallback_sortedness_retracted, any other unusable build keeps fallback_metallib_unavailable, then re-routes to the legacy path as before. F4: new doctest drives the full armed() -> clear_and_arm() -> snapshot_and_disarm() cycle and the attempts == hits + fallbacks invariant including the new class; the route-table and counter-invariant tests now cover fallback_sortedness_retracted. Verified: cmake tests 262/262 + 3550 assertions pass; metal -Wall -Wextra -fno-fast-math compile of kernels/quantized.metal is warning-free. * perf(metal): E=256 expert-tile route + trust + gpu::eval UAF fix — darkbloom-base mirror (#7) * perf(metal): instantiate E=256 expert-tile route for Qwen 3.5/3.6 MoE prefill (mirror of Cmlx/mlx 58fab46) * fix(metal): use-after-free in gpu::eval for primitives that synchronize mid-eval (mirror) * perf(metal): trust mode skips retract readback (mirror) * fix(compile): preserve all-cache binding cleanup --------- Co-authored-by: JasonHonKL <148705846+JasonHonKL@users.noreply.github.com> Co-authored-by: AK <144495202+AKnassa@users.noreply.github.com> Co-authored-by: Cheng <git@zcbenz.com> Co-authored-by: Adityaj0 <93090622+Adityaj0@users.noreply.github.com> Co-authored-by: anchor <codeanqiang@gmail.com> Co-authored-by: codeAnqiang-ma <273298913+codeAnqiang-ma@users.noreply.github.com> Co-authored-by: Rohan Gautam <rohan1gautam@gmail.com> Co-authored-by: Ayaan Gazali <ayaangazali.work@gmail.com> Co-authored-by: Erwin Zhang <59893706+erwinzhang7@users.noreply.github.com> Co-authored-by: katlun-lgtm <katlun@gmail.com> Co-authored-by: katlun-lgtm <264247399+katlun-lgtm@users.noreply.github.com> Co-authored-by: robertomeroni <150194833+robertomeroni@users.noreply.github.com> Co-authored-by: Fu Xiaonan <ht3fudatou@163.com> Co-authored-by: Fu Xiaonan <214359569+FU-max-boop@users.noreply.github.com> Co-authored-by: Hao Xu <hxu44@apple.com> Co-authored-by: Feli <89400571+FeliGame@users.noreply.github.com> Co-authored-by: Feli <feli@hnu.edu.cn> Co-authored-by: Eyüp Can Akman <eyupcanakman@gmail.com> Co-authored-by: Cheng <zcbenz@gmail.com> Co-authored-by: yentur <mr.yentur@gmail.com> Co-authored-by: Alessio Pollero <alessio.pollero@gmail.com> Co-authored-by: Zhiqi Zhang <zhiqizhangg@gmail.com> Co-authored-by: Daniel Hiltgen <dhiltgen@users.noreply.github.com> Co-authored-by: Tanish Jain <recklurker@gmail.com> Co-authored-by: hojin12312 <hojin12312@gmail.com> Co-authored-by: Xiang Chen <46052474+x14ngch3n@users.noreply.github.com> Co-authored-by: x14ngch3n <x14ngch3n@users.noreply.github.com> Co-authored-by: Duhyeon, Kim <49020301+dudududukim@users.noreply.github.com> Co-authored-by: rohith <kapellirohith@gmail.com> Co-authored-by: Ishaan Samantray <devteam.aegis@gmail.com> Co-authored-by: Aaishwarya Mishra <aaishwarymishra@gmail.com> Co-authored-by: Anastasiia Filippova <a_filippova@apple.com> Co-authored-by: Yanzhao Wang <19340816+wyanzhao@users.noreply.github.com> Co-authored-by: XXXXRT666 <157766680+XXXXRT666@users.noreply.github.com> Co-authored-by: Jake Bowhay <60778417+j-bowhay@users.noreply.github.com> Co-authored-by: vraj patel <87225460+vraj00222@users.noreply.github.com> Co-authored-by: Gusanidas <33495733+Gusanidas@users.noreply.github.com> Co-authored-by: Dwijen Patel <dwijen@gmail.com> Co-authored-by: Vladimir Iglovikov <ternaus@users.noreply.github.com> Co-authored-by: Brian C. <94733710+deBrian07@users.noreply.github.com> Co-authored-by: Daniel Hiltgen <daniel.hiltgen@ollama.com> Co-authored-by: YH Yan <strayberry0w0@gmail.com> Co-authored-by: katlun-lgtm <katlun@windyviews.com> Co-authored-by: anupsv <6407789+anupsv@users.noreply.github.com> Co-authored-by: Gajesh Naik <26431906+Gajesh2007@users.noreply.github.com> Co-authored-by: David Tai <davidtai@Davids-MBP.lan>
…0119 releases, whose conv3d (kT=1) routes mlx-swift 0.32 defeats R-NUM-1 (d47561a) justified leaving TF32 on by saying packages already fix the Winograd window locally, and pointed at the 14 image releases of AB-A-0119 (2026-09-25). mlx-forge's close of that ask, and the 0.32 wave's migration table (mlxengine-todo/wave-0.32-approval-batch.md §B), show otherwise: 13 of those routes go through conv3d with kT = 1, which mlx-swift 0.32 splits back into the Winograd conv2d (ml-explore/mlx#3785). gfpgan's route measured 1.01e-3 on 0.32.3. The fix that holds on 0.31.x and 0.32.x is MLXExactConv (mlx-exact-conv-swift ≥ 0.1.0), from the 2026-10-01 migration. The decision is unchanged: TF32 stays on; the per-package exact route is now named correctly, with a dated correction note (the AB-L-0111 failure mode: a doc pointing at a fix that a later version undid). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Closes #3625.
Problem
mx.conv_generalwith 5D inputs (3D convolution) is 2–5× slower than decomposing thesame op into per-frame 2D convolutions with a Python loop. This hits video-generation
workloads (VAE decoders with
CausalConv3d).The reason isn't the explicit-gemm fallback — for the common case (mod16 channels,
idil==1,groups==1)dispatch_conv_3D_gpualready routes toimplicit_gemm_conv_3D_gpu. The real cause is that the 2D dispatch has a Winogradkernel for 3×3 stride-1 convs (
dispatch_conv_2D_gpu→winograd_conv_2D_gpu), and the3D path has no Winograd or 3×3-specialized kernel — so a 3×3×3 conv misses Winograd
entirely. Decomposing a small kernel-depth 3D conv into
KD2D convs lets each 2D convhit the tuned 2D path.
Change
small_kd_conv_3D_gpu: for small kernel depth, for each depth tapkdwe build azero-copy strided view of the input frames (
[OD, H, W, C]at depth offsetkd) and theweight depth-slice (
[O, KH, KW, C]), runconv_2D_gpu, and accumulate. The accumulatorbuffer is repointed into
outviacopy_shared_buffer(thepad_and_slicepattern).A guard in
dispatch_conv_3D_gputakes this path only when it is valid and faster:idil == 1(all dims),groups == 1,N == 1;KD <= 7(theKD-1accumulate adds erode the win for largeKD);stride == 1andkernel_dilation == 1, andpad[0] == 0;mod16channels (so the per-frame 2D convs hit the fast path).Everything else falls through to the existing implicit gemm unchanged.
(Skeleton follows @Ved235's sketch in the issue; the depth-tap decomposition idea is
theirs.)
Results (M3 Max, mlx built from this branch)
Correctness — the fast path vs the CPU reference (fp32, exact):
Speed (bf16, 3×3×3), native 3D vs the per-frame 2D decomposition:
i.e. 2.4–3.4× → ~1.0× (parity) with the per-frame 2D path, and numerically identical
to it.
Tests
python/tests/test_conv.py:test_conv_3D_small_kd_decomposition— the fast path vs CPU across 5 shapes(
Cout != Cin,KD ∈ {1,2,3,5}, 1×1 spatial), fp32 exact.test_conv_3D_small_kd_fallback_cases— depth stride > 1, depth padding, and non-mod16channels stay correct (must take the fall-back path).
Full
test_conv.pysuite passes with no regressions.Open questions for maintainers
binary_op_gpu_inplace(..., "Add"). Prefer that, or abeta-accumulate added toconv_2D_gpu?KDthreshold — fixed at 7, or a cost heuristic vs the 3D implicit gemm?later? Also: should the fast path be extended to
N > 1and depth padding, or is theN == 1/pad[0] == 0guard fine for a first PR?