Conversation
|
@jhen0409 are the perf improvements measurable? |
81e2817 to
260eebc
Compare
Yeah it's measurable. Rebased. TG has only slight gains of 0-10% after #29199. (updated PR description) GGML_HEXAGON_PROFILE=3 GGML_HEXAGON_OPTRACE=8192 GGML_HEXAGON_OPBATCH=1 \
test-backend-ops test -b HTP0:0 -o FLASH_ATTN_EXT --test-file cases.txt > trace.log 2>&1
python3 scripts/snapdragon/ggml-hexagon-profile.py trace.log --timeline summary --filter FLASH_ATTNcases.txt (p-d64-h32-kh8-k4096-nomask-s1024 is reference): master (after #29199)PR (rebased) |
|
Here is a better version #29282 |
Yeah that's better. Thanks! |
Overview
This PR is an improvement for
flash_attn_ext_f16_threadwhich replaces the original cache search on every push.The FA mask blocks are pushed in the same order for every head: (token, block) within 1 mask head. So sizing the cache to that cycle (
n_blocks * neq1 * mask->ne[2]) (<= 128) is enough. Thatdma_cache_push_seq(new function) can find every repeat without a search.Additional information
Setup: IQ-9075 (Hexagon v73), single NPU HTP0:0,
llama-bench -p 512 -n 64 -r 2. master is 260eebc.Old (before #29199)
PPL not changed and pp512 t/s keep original.
Requirements