Skip to content

Implement QK attention head chunking in MLA to reduce HBM footprint#4564

Merged
copybara-service[bot] merged 1 commit into
mainfrom
zjiahao/DSA3.2-qk-chunk
Jul 24, 2026
Merged

Implement QK attention head chunking in MLA to reduce HBM footprint#4564
copybara-service[bot] merged 1 commit into
mainfrom
zjiahao/DSA3.2-qk-chunk

Commits

Commits on Jul 24, 2026