forked from vllm-project/vllm
-
Notifications
You must be signed in to change notification settings - Fork 8
Pull requests: MooreThreads/vllm-musa
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
MUSA: auto-dispatch FP8 MoE GEMV and GEMM on S5000
#114
opened Jul 21, 2026 by
yeahdongcn
Collaborator
•
Draft
9 tasks done
perf: add fast top-k renormalization on MUSA
#112
opened Jul 20, 2026 by
yeahdongcn
Collaborator
•
Draft
perf: add chunked min-p sampling on MUSA
#111
opened Jul 20, 2026 by
yeahdongcn
Collaborator
•
Draft
perf: fuse residual RMSNorm and FP8 group quant
#110
opened Jul 20, 2026 by
yeahdongcn
Collaborator
•
Draft
perf: fuse SiLU and FP8 group quant before DeepGEMM
#109
opened Jul 20, 2026 by
yeahdongcn
Collaborator
•
Draft
Coerce hybrid-model speculative decode off full decode CUDAGraphs on MUSA
#97
opened Jul 12, 2026 by
yeahdongcn
Collaborator
•
Draft
5 tasks done
Fuse Qwen3 dense-decode q-norm + RoPE into one kernel (opt-in)
#90
opened Jul 6, 2026 by
yeahdongcn
Collaborator
•
Draft
6 tasks done
ProTip!
Follow long discussions with comments:>50.