Skip to content

[pull] master from ggml-org:master#205

Merged
pull[bot] merged 2 commits intoLongLeCE:masterfrom
ggml-org:master
Jul 27, 2025
Merged

[pull] master from ggml-org:master#205
pull[bot] merged 2 commits intoLongLeCE:masterfrom
ggml-org:master

Conversation

@pull
Copy link
Copy Markdown

@pull pull Bot commented Jul 27, 2025

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.3)

Can you help keep this open source service alive? 💖 Please sponsor : )

deepsek and others added 2 commits July 27, 2025 00:28
…14624)

This commit adds support for MFMA instructions to MMQ. CDNA1/GFX908 CDNA2/GFX90a and CDNA3/GFX942 are supported by the MFMA-enabled code path added by this commit. The code path and stream-k is only enabled on CDNA3 for now as it fails to outperform blas in all cases on the other devices.
Blas is currently only consistently outperformed on CDNA3 due to issues in the amd-provided blas libraries.
This commit also improves the awareness of MMQ towards different warp sizes and as a side effect improves the performance of all quant formats besides q4_0 and q4_1, which regress slightly, on GCN gpus.
@pull pull Bot locked and limited conversation to collaborators Jul 27, 2025
@pull pull Bot added the ⤵️ pull label Jul 27, 2025
@pull pull Bot merged commit 446595b into LongLeCE:master Jul 27, 2025
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants