forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 10
Pull requests: Nathanw1014/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
vulkan: sparse prefill flash attention for qwen4exp top-k masks
CUDA
documentation
Improvements or additions to documentation
ggml
model
testing
Vulkan
qwen4exp: Qwen3.8-Flash-Next long-context decode, +66% at 131k
model
#9
opened Sep 7, 2026 by
firelzrd
Loading…
spec: adaptive draft sizing (--spec-draft-adaptive) - port of LaurentZuijdwijk's controller
#8
opened Sep 2, 2026 by
aic0d3r
Loading…
ProTip!
Adding no:label will show everything without a label.