You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: _posts/2025-07-31-cachegen.md
+21-3Lines changed: 21 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -3,11 +3,11 @@ layout: post
3
3
title: "CacheGen: Store Your KV Cache on Disk or S3—Load Blazingly Fast!"
4
4
thumbnail-img: /assets/img/cachegen.png
5
5
share-img: /assets/img/cachegen.png
6
-
author: Kuntai Du
6
+
author: Kuntai Du, Kobe
7
7
image: /assets/img/cachegen.png
8
8
---
9
9
10
-
**TL;DR:** 🚀 CacheGen lets you store KV caches on disk or AWS S3 and load them *way* faster than recomputing! It compresses your KV cache up to **3× smaller than quantization** so that you can load your KV cache blazingly fast while keeping response quality high. Stop wasting compute --- use CacheGen to fully utilize your storage and get instant first-token speedup!
10
+
**TL;DR:** 🚀 [CacheGen](https://arxiv.org/abs/2310.07240) lets you store KV caches on disk or AWS S3 and load them *way* faster than recomputing! It compresses your KV cache up to **3× smaller than quantization** so that you can load your KV cache blazingly fast while keeping response quality high. Stop wasting compute --- use CacheGen to fully utilize your storage and get instant first-token speedup!
If you use CacheGen in your research, please cite our paper:
75
+
76
+
```bibtex
77
+
@misc{liu2024cachegenkvcachecompression,
78
+
title={CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving},
79
+
author={Yuhan Liu and Hanchen Li and Yihua Cheng and Siddhant Ray and Yuyang Huang and Qizheng Zhang and Kuntai Du and Jiayi Yao and Shan Lu and Ganesh Ananthanarayanan and Michael Maire and Henry Hoffmann and Ari Holtzman and Junchen Jiang},
80
+
year={2024},
81
+
eprint={2310.07240},
82
+
archivePrefix={arXiv},
83
+
primaryClass={cs.NI},
84
+
url={https://arxiv.org/abs/2310.07240},
85
+
}
86
+
```
87
+
88
+
**Paper:**[CacheGen: KV Cache Compression and Streaming for Fast Large Language Model Serving](https://arxiv.org/abs/2310.07240)
0 commit comments