Skip to content

Add oneliner for batch quantization - #17

Closed
jooray wants to merge 1 commit into
ggml-org:masterfrom
jooray:patch-2
Closed

jooray wants to merge 1 commit into
ggml-org:masterfrom
jooray:patch-2

Conversation

@jooray

@jooray jooray commented Mar 11, 2023

Copy link
Copy Markdown
Contributor

No description provided.

Comment thread README.md
which can be scripted like this if you are lazy (for 65B model):

```bash
for i in models/65B/ggml-model-f16.bin*;do quantized=`echo "$i" | sed -e 's/f16/q4_0/'`; ./quantize "$i" "$quantized" 2 ;done

@prusnak prusnak Mar 11, 2023 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sed is not necessary, bash, zsh and other modern shells can perform pattern replacement of a variable:

Suggested change
for i in models/65B/ggml-model-f16.bin*;do quantized=`echo "$i" | sed -e 's/f16/q4_0/'`; ./quantize "$i" "$quantized" 2 ;done
for i in models/65B/ggml-model-f16.bin* ; do ./quantize "$i" "${i/f16/q4_0}" 2 ;done

@s-and-witch s-and-witch Mar 12, 2023 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This will generate 'models/65B/ggml-model-q4_0/.bin.2' such paths and will fail with errors, the right command (in bash) should be for i in models/65B/ggml-model-f16.bin* ; do ./quantize "$i" "${i/f16/q4_0}" 2 ;done

@prusnak prusnak Mar 12, 2023 •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@Player-205 right, updated the suggestion above, thanks

@ggerganov

Copy link
Copy Markdown
Member

Lets put this in a quantize.sh script that accepts argument like 7B, 13B, etc. and update instructions to just run the script:

source quantize.sh 7B

Should be much easier to follow

@leszekhanusz

Copy link
Copy Markdown

Note that if the disk space is limited, it is still useful to quantize each file separately so that we could delete each intermediate file in between.
In my case I added a rm command because I did not have enough disk space otherwise:

for i in models/65B/ggml-model-f16.bin* ; do ./quantize "$i" "${i/f16/q4_0}" 2 ; rm "$i"; done

@ggerganov

Copy link
Copy Markdown
Member

Good point, should have a second parameter for "keep f16" which is on by default

@prusnak

prusnak commented Mar 13, 2023

Copy link
Copy Markdown
Contributor

Superseded by #92

@ggerganov ggerganov closed this Mar 13, 2023
SlyEcho pushed a commit to SlyEcho/llama.cpp that referenced this pull request Jun 11, 2023
jesusmb1995 pushed a commit to jesusmb1995/llama.cpp that referenced this pull request Sep 29, 2025
QVAC-5545: Use char instead uint8_t for streams
wine99 pushed a commit to wine99/llama.cpp that referenced this pull request Nov 27, 2025
This was referenced Nov 28, 2025
@jacekpoplawski jacekpoplawski mentioned this pull request Feb 10, 2026
1 of 3 tasks
reeselevine added a commit that referenced this pull request Feb 18, 2026
* Basic JIT compilation for mul_mat, get_rows, and scale (#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* flashattention and matrix multiplication moved to new format

* clean up preprocessing

* Formatting

* remove duplicate constants

* Split large shaders into multiple static strings

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
liparetejas pushed a commit to liparetejas/llama.cpp that referenced this pull request Feb 23, 2026
* Basic JIT compilation for mul_mat, get_rows, and scale (ggml-org#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* flashattention and matrix multiplication moved to new format

* clean up preprocessing

* Formatting

* remove duplicate constants

* Split large shaders into multiple static strings

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
reeselevine added a commit that referenced this pull request Mar 10, 2026
…better shader parameter handling (#20173)

* K quant speedup (#20)

* Basic JIT compilation for mul_mat, get_rows, and scale (#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* no gibberish, all k quants added, merged

* vec memory fix

* q6_k matching metal on my machine, tests passing

* Set tile size for q6_k separately

* Separate out fast shaders

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>

* Move towards writeBuffer for params

* Move away from multiple buffers for set_rows errors, remove host buffer for parameter buffers, minor cleanups

* Remove extra file

* Formatting

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
platima pushed a commit to platima/llama.cpp-SpacemiT-K3 that referenced this pull request Jul 9, 2026
…-only)

Cherry-pick of the docs portion of c0a31b5 only: adds
docs/spacemit-mtmd-model-development.md (HF VLM -> GGUF text + ONNX vision split
and SMT pipeline development) and links it from docs/spacemit-mtmd.md.

The OHOS build-slimming and CI-disable hunks from the original commit are dropped
here — this K3 fork does not carry the OpenHarmony (OHOS) target (ggml-org#16 skipped).

(cherry picked from commit c0a31b5, docs hunks only)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* Basic JIT compilation for mul_mat, get_rows, and scale (#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* flashattention and matrix multiplication moved to new format

* clean up preprocessing

* Formatting

* remove duplicate constants

* Split large shaders into multiple static strings

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
…better shader parameter handling (#20173)

* K quant speedup (#20)

* Basic JIT compilation for mul_mat, get_rows, and scale (#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* no gibberish, all k quants added, merged

* vec memory fix

* q6_k matching metal on my machine, tests passing

* Set tile size for q6_k separately

* Separate out fast shaders

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>

* Move towards writeBuffer for params

* Move away from multiple buffers for set_rows errors, remove host buffer for parameter buffers, minor cleanups

* Remove extra file

* Formatting

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* ggml: backend-agnostic tensor parallelism

* support for GPT-OSS, Qwen 3 MoE

* partial Vulkan fix

* add support for 4/8 GPUs

* unconditional peer access

* re-use buffers + ggml contexts

* fix output pattern

* NCCL support

* GGML: HIP: add RCCL support

* Remove shfl and AllReduce from backend interface

* move allocation workaround out of ggml-alloc.c

* 2d tensor set/get support

* Fix the seg fault without NCCL

* Apply suggestion from JohannesGaessler

* support for tensor dims % n_devs != 0

* fix view_offs scaling

* arbitrary num. of GPUs/tensor split

* fix compilation

* better granularity estimate

* Support device-specific host buffer types if all underlying backends expose the same type. This allows using pinned memory instead of pageable memory for CUDA.

Fix compilation errors.

* partial Qwen 3 Next support

* Fix qwen3 30b (#8)

* Fix crash with Qwen-30B-A3B Q4_0

Qwen-30B-A3B Q4_0 has an intermediate dimension of 768. Using a granularity of 256 forces an uneven split between GPUs, which is not supported by the current implementation.

* Decide block size based on tensor quantization type

* Fix crashes due to KV cache serialization (#9)

KV cache serialization requires non-zero offsets on the tensor. Add support in the meta backend to set/get a tensor with a non-zero offset.

* metal : fix build (#7)

* static memory allocations, fix usage count

* fix tensor granularity

* more even memory distribution

* use BF16 for allreduce

* rebase fixup

* better error message for unsupported architectures

* Fix device mismatch during scatter of allReduce. (#11)

There is a mismatch between the dst buffer device and the backend device, causing the use of sync copies

* Enable the previous allreduce implementation. It is better in both perf and stability (#12)

* delay AllReduce for Moe for less I/O

* build : clean-up compile warnings

* backend : move most of the meta backend API to ggml-backend-impl.h

* cont : hide unused public API in the implementation

* llama : use llama_device + remove ggml_backend_dev_is_meta()

* ggml-backend : remove unused alloc include

* minor : remove regex include

* ggml : introduce ggml-ext.h for staging new APIs

* rebase fixup

* fix tests

* llama : more robust logic for determining Meta devices (#16)

* llama : more robust logic for determining Meta devices

* cont : fix devs size check

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* cont : fix log type

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* disable roundtrip for meta backend

* fix arch selection

* Qwen 3.5 support

* fix Gemma 4 MoE

* fix OpenVino, SYCL

* fix test-llama-archs for CPU-only builds

* Fix Qwen 3.5 MoE

* disable meta backend tests for WebGPU

* tests : filter CPU-based devices from the Meta backend tests (#17)

* meta : formatting, naming, indentation (#18)

* formatting : llama-model.cpp

* formatting : ggml-ext.h

* formatting : ggml-backend-meta.cpp

* meta : add TODO

* add documentation

* better error messages

* fix GPT-OSS

---------

Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
Co-authored-by: Gaurav Garg <gaugarg@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
juliantorr-es pushed a commit to Tribunus-dev/tessera that referenced this pull request Aug 14, 2026
Second, deliberately-scoped-down half of the RecalcScheduler bundle
(0.2a shipped RecalcState earlier this session; this is the "TokenArray
IR + shared formula groups" the plan's capability ggml-org#17 called for -
"covered (RPN)", matching LibreOffice's own internal representation).

TokenArray.swift: a flat RPN token stream compiled from a FormulaAST,
with cell/range references encoded RELATIVE TO THE FORMULA'S OWN
ANCHOR CELL rather than as absolute addresses - the property that
makes two formulas "the same shape" (=A1+B1 at B2 and =A2+B2 at B3
both compile to the identical stream), which is what lets
SharedFormulaGroups intern one compiled representation per shape
instead of allocating a new one for every cell a fill-down produces.
RPNToken (named to avoid colliding with Lexer's own unrelated Token
type - hit that ambiguity immediately on the first build), a compiler
(FormulaAST + anchor -> RPNToken stream, returning nil for LET/LAMBDA
and array literals - documented, bounded scope, not silently gapped),
and a stack-machine evaluator.

Extracted BinaryEvaluator/UnaryEvaluator (TypeSystem.swift) out of
Evaluator.swift's private evalBinary/evalUnary bodies - pure functions
of already-evaluated Values, no cell/sheet/environment context needed,
so both the existing AST-walking Evaluator and the new
TokenArrayEvaluator call the SAME logic instead of risking two copies
silently diverging on what "+" means. Evaluator.swift now delegates
to both instead of duplicating them - behavior-preserving refactor,
covered by the full existing test suite.

Deliberately NOT wired into SheetEngine's live cell-evaluation path -
see the file header's scope note. Swapping the production evaluation
path is separate, additional risk (correctness-critical, hot-path
code) beyond "add a new, optional, tested capability." FormulaAST
stays the surface SheetEngine actually evaluates; TokenArray is ready
infrastructure for a later increment to adopt.

Verified via a real `swift test` run (temporary, never-committed
neutralization of the same three unrelated pre-existing Data-layer
breaks used throughout this session, reverted after): the load-bearing
property - TokenArrayEvaluator must produce EXACTLY what the existing,
production Evaluator produces for the same formula/cells - is checked
directly in TokenArrayTests.swift (both evaluators run against the
same live SheetEngine, not asserted from reading the code), across
arithmetic, comparisons, cell/range refs, absolute vs. relative mixes,
nested functions, named ranges, cross-sheet references, and error
propagation. Also covers what does NOT compile (LET/LAMBDA/array
literals), that same-shape-different-anchor formulas produce equal
group keys (the actual "shared formula groups" property), and
SharedFormulaGroups' intern/release/refcount lifecycle.
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* Basic JIT compilation for mul_mat, get_rows, and scale (ggml-org#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* flashattention and matrix multiplication moved to new format

* clean up preprocessing

* Formatting

* remove duplicate constants

* Split large shaders into multiple static strings

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
…better shader parameter handling (ggml-org#20173)

* K quant speedup (ggml-org#20)

* Basic JIT compilation for mul_mat, get_rows, and scale (ggml-org#17)

* scale jit working

* preliminary working jit for getrows and mulmat, needs refining

* simplified mul_mat preprocessing switch statement

* get_rows fixes, mul_mat refinement

* formatted + last edits

* removed some extraneous prints

* fixed get_rows, fixed workgroup dispatch in mul_mat. no gibberish

* small fix

* some changes, working

* get_rows and mul_mat jit fixed and working

* Update formatting

* formatting

* Add header

---------

Co-authored-by: Neha Abbas <nehaabbas@ReeseLevines-MacBook-Pro.local>
Co-authored-by: Reese Levine <reeselevine1@gmail.com>

* Start work on all-encompassing shader library

* refactor argmax, set_rows

* Refactor all but flashattention, mat mul

* no gibberish, all k quants added, merged

* vec memory fix

* q6_k matching metal on my machine, tests passing

* Set tile size for q6_k separately

* Separate out fast shaders

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>

* Move towards writeBuffer for params

* Move away from multiple buffers for set_rows errors, remove host buffer for parameter buffers, minor cleanups

* Remove extra file

* Formatting

---------

Co-authored-by: neha-ha <137219201+neha-ha@users.noreply.github.com>
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
)

* ggml: backend-agnostic tensor parallelism

* support for GPT-OSS, Qwen 3 MoE

* partial Vulkan fix

* add support for 4/8 GPUs

* unconditional peer access

* re-use buffers + ggml contexts

* fix output pattern

* NCCL support

* GGML: HIP: add RCCL support

* Remove shfl and AllReduce from backend interface

* move allocation workaround out of ggml-alloc.c

* 2d tensor set/get support

* Fix the seg fault without NCCL

* Apply suggestion from JohannesGaessler

* support for tensor dims % n_devs != 0

* fix view_offs scaling

* arbitrary num. of GPUs/tensor split

* fix compilation

* better granularity estimate

* Support device-specific host buffer types if all underlying backends expose the same type. This allows using pinned memory instead of pageable memory for CUDA.

Fix compilation errors.

* partial Qwen 3 Next support

* Fix qwen3 30b (ggml-org#8)

* Fix crash with Qwen-30B-A3B Q4_0

Qwen-30B-A3B Q4_0 has an intermediate dimension of 768. Using a granularity of 256 forces an uneven split between GPUs, which is not supported by the current implementation.

* Decide block size based on tensor quantization type

* Fix crashes due to KV cache serialization (ggml-org#9)

KV cache serialization requires non-zero offsets on the tensor. Add support in the meta backend to set/get a tensor with a non-zero offset.

* metal : fix build (ggml-org#7)

* static memory allocations, fix usage count

* fix tensor granularity

* more even memory distribution

* use BF16 for allreduce

* rebase fixup

* better error message for unsupported architectures

* Fix device mismatch during scatter of allReduce. (ggml-org#11)

There is a mismatch between the dst buffer device and the backend device, causing the use of sync copies

* Enable the previous allreduce implementation. It is better in both perf and stability (ggml-org#12)

* delay AllReduce for Moe for less I/O

* build : clean-up compile warnings

* backend : move most of the meta backend API to ggml-backend-impl.h

* cont : hide unused public API in the implementation

* llama : use llama_device + remove ggml_backend_dev_is_meta()

* ggml-backend : remove unused alloc include

* minor : remove regex include

* ggml : introduce ggml-ext.h for staging new APIs

* rebase fixup

* fix tests

* llama : more robust logic for determining Meta devices (ggml-org#16)

* llama : more robust logic for determining Meta devices

* cont : fix devs size check

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* cont : fix log type

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* disable roundtrip for meta backend

* fix arch selection

* Qwen 3.5 support

* fix Gemma 4 MoE

* fix OpenVino, SYCL

* fix test-llama-archs for CPU-only builds

* Fix Qwen 3.5 MoE

* disable meta backend tests for WebGPU

* tests : filter CPU-based devices from the Meta backend tests (ggml-org#17)

* meta : formatting, naming, indentation (ggml-org#18)

* formatting : llama-model.cpp

* formatting : ggml-ext.h

* formatting : ggml-backend-meta.cpp

* meta : add TODO

* add documentation

* better error messages

* fix GPT-OSS

---------

Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz>
Co-authored-by: Gaurav Garg <gaugarg@nvidia.com>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
Tha14 pushed a commit to Tha14/llama.cpp-wackMall-merge-request that referenced this pull request Sep 6, 2026
Critical bug: GGML_OP_SYMBOL was missing entries for GATED_DELTA_NET_TREE
and SSM_CONV_TREE, causing every symbol from op 86 onward to be shifted
2 positions. This is the class of bug that produces 'op not implemented: 98'
crashes in shipped binaries (ggml-org#17). Fixed by converting both GGML_OP_NAME
and GGML_OP_SYMBOL to designated initializers and adding assertions to
ggml_op_name()/ggml_op_symbol() so any future drift is caught at runtime.

CUDA kernel hardening:
- Add bounds guards (h_idx >= H || sequence >= n_seqs || col >= S_v) to
  both gated_delta_net_cuda and gated_delta_net_tree_cuda kernels, matching
  the existing dflash_gdn_state_replay_cuda guard. Padded grid dimensions
  from ceil division can produce out-of-bounds col values.
- Add full shape/type assertions to ggml_cuda_op_gated_delta_net (non-tree)
  matching the tree function's existing validation. Fails early with clear
  messages instead of silent CUDA memory corruption.

DFlash replay stream ordering (ggml-org#15 root cause):
- Add dflash_replay_gdn_state_with_stream() that accepts an explicit
  cudaStream_t instead of using cudaStreamPerThread. The old entry point
  becomes a compatibility wrapper passing nullptr (falls back to per-thread
  stream). Correct stream ordering prevents races between backend stream
  work and replay kernel launches, which surfaced as illegal memory access
  at cudaStreamSynchronize.
- Add dflash_cuda_ptr_device_visible() validation that rejects host-memory
  pointers before launching direct CUDA replay. Returns false (allowing CPU
  fallback) instead of silently corrupting device memory.
- Expose ggml_backend_cuda_get_stream() API and register it in proc_address
  so llama-context can pass the correct backend stream to replay.
- Update both tape_replay_gdn_direct_gpu and
  tape_replay_gdn_direct_from_cpu_tape to prefer the stream-taking replay
  function, falling back to the old entry point if unavailable.

Op coverage regression test:
- Add tests/test-ggml-op-coverage.cpp that iterates all GGML_OP_COUNT ops
  and verifies every op has a non-empty name and symbol. Catches enum/table
  drift before release.

Build stamps:
- Add git dirty detection to cmake/build-info.cmake (BUILD_DIRTY flag).
- Add dirty suffix to llama_build_info() and ggml op count to
  llama_print_build_info() so users can identify stale binaries.
Tha14 pushed a commit to Tha14/llama.cpp-wackMall-merge-request that referenced this pull request Sep 6, 2026
Remove ggml-org#15-only stream-safe direct replay code that introduced
dflash_capture->backend references and broke the Windows build:

- Revert gated_delta_net.cu: remove dflash_replay_gdn_state_with_stream,
  dflash_cuda_ptr_device_visible, bounds guards, and shape assertions
  (all ggml-org#15 hardening, not ggml-org#17)
- Revert ggml-cuda.h: remove ggml_backend_cuda_get_stream declaration
- Revert ggml-cuda.cu: remove stream proc_address entries and
  ggml_backend_cuda_get_stream implementation
- Revert llama-context.cpp: remove stream-taking replay calls and
  dflash_capture->backend references
- Revert test-dflash-plumbing.cpp: remove stream/pointer validation checks

The ggml-org#17 op coverage fix remains:
- designated initializers for GGML_OP_NAME/GGML_OP_SYMBOL
- missing GATED_DELTA_NET_TREE and SSM_CONV_TREE symbol entries
- op coverage test
- build stamp dirty flag and GGML_OP_COUNT
SimonTeixidor pushed a commit to SimonTeixidor/llama.cpp that referenced this pull request Sep 13, 2026
Strix Halo Vulkan stack: FA KV-quant, mmid, dense GEMM, delta-net prefill, DSv4 sparse attention
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants