forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 8
Pull requests: ROCm/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
mmq: opt-in compacted MoE tiling for RDNA3.5
#63
opened Jul 21, 2026 by
roberteg16
•
Draft
2 of 3 tasks
tests: add MoE MMQ benchmark with routing-distribution generator
#62
opened Jul 20, 2026 by
roberteg16
•
Draft
3 tasks
ggml-cuda: add dequant-float matvec (mmvdq) for Q4_K/Q5_K/Q6_K on RDN…
#61
opened Jul 20, 2026 by
Annieren
Loading…
gfx1151 (Strix Halo): fuse attn_k+v into single MMVQ dispatch
#59
opened Jul 17, 2026 by
jeffli-xilinx
Loading…
7 tasks done
ggml-cuda: GEMM weight row padding + one-time K-padded f16 dequant for prefill
#57
opened Jul 17, 2026 by
roberteg16
•
Draft
4 of 5 tasks
CUDA: gated_delta_net - share per-token k/q via shared memory for long prefills
#54
opened Jul 16, 2026 by
roberteg16
•
Draft
1 of 2 tasks
RDNA3.5 (gfx11 / gfx1151) MMQ prefill optimizations
#32
opened Jun 30, 2026 by
liangliangchang
Loading…
4 tasks done
ProTip!
Add no:assignee to see everything that’s not assigned.