Skip to content

Commit 48209e8

Browse files
committed
update PR number
1 parent acf96ac commit 48209e8

1 file changed

Lines changed: 1 addition & 1 deletion

File tree

perf-changelog.yaml

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -4929,4 +4929,4 @@
49294929
- "Add GB200 Dynamo-vLLM AgentX aggregate TP8 at conc [1,4,8,16] and disaggregated topologies: 1P/1D DEP8/DEP8 at [64,128,256,320], 1P/1D DEP8/DEP12 at [128,256,320], 2P/1D DEP8/DEP12 at [480,640,768], and 3P/1D DEP8/DEP16 at [640,800,960,1280]."
49304930
- "Aggregate TP8 uses sparse DSV4 attention, NUMA binding, max-num-seqs/CUDA graph size 32, max-num-batched-tokens 32768, and 0.95 GPU memory utilization."
49314931
- "Image: vllm/vllm-openai:nightly-dev-arm64-cu13.0.1-c188b96"
4932-
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/YYYY
4932+
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2260

0 commit comments

Comments
 (0)