Skip to content

Implement QK attention head chunking in MLA to reduce HBM footprint#4564

Merged
copybara-service[bot] merged 1 commit into
mainfrom
zjiahao/DSA3.2-qk-chunk
Jul 24, 2026
Merged

Implement QK attention head chunking in MLA to reduce HBM footprint#4564
copybara-service[bot] merged 1 commit into
mainfrom
zjiahao/DSA3.2-qk-chunk

[Memory] Implement QK attention head chunking to prevent HBM OOM

3f588a0
Select commit
Loading
Failed to load commit list.
Google CLA / cla/google succeeded Jul 24, 2026 in 7s

✅ All contributors are covered under a CLA with Google

See https://cla.developers.google.com/ for more info about Google's Contributor License Agreement (CLA).

ℹ️ Googlers: Go here to view more details and manage scans for this pull request.

Details

The following contributors were found for this pull request:

3f588a0 Author: @zcjhao <zj****o​@google.com>

(Only the first commit for a unique contributor is listed.)