[NV] add dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k / 新增 DeepSeek-V4-Pro FP4 GB300 Dynamo SGLang MTP 1k1k 基准测试配置#1697
[NV] add dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k / 新增 DeepSeek-V4-Pro FP4 GB300 Dynamo SGLang MTP 1k1k 基准测试配置#1697hshrivastava-droid wants to merge 32 commits into
Conversation
|
Thanks for the contribution! For vLLM & SGLang, please ensure that your recipes is similar to the official vLLM recipes and/or the SGLang cookbook If it is not, please create a PR first before we can merge your single node PR into the master branch. Let's ensure that the documentation is first class such that the entire ML community can benefit from your hard work! Thank you PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. If re-running failed jobs is attempted, PR authors are responsible for ensuring it passes. See GitHub's docs on re-running failed jobs: https://docs.github.com/en/actions/how-tos/manage-workflow-runs/re-run-workflows-and-jobs#re-running-failed-jobs-in-a-workflow As a rule of thumb, generally, PR authors should request a review & get a PR approval from the respective companies' CODEOWNERS before requesting a review from core maintainers. If additional help is needed, PR authors can reach out to core maintainers over Slack. |
|
|
||
| model: | ||
| path: "dsv4-pro" | ||
| container: "lmsysorg/sglang:nightly-dev-cu13-20260510-2473659e" |
There was a problem hiding this comment.
Low-latency recipe container missing
High Severity
The two low-latency recipes still pin model.container to lmsysorg/sglang:nightly-dev-cu13-20260510-2473659e, while dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k imports squash only for lmsysorg/sglang:nightly-dev-cu13-20260603-83bc7766. Workers resolve the recipe tag, which is not mapped in srtslurm.yaml and is documented as absent from Docker Hub, so those matrix points can fail at enroot import.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 2aeafb4. Configure here.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27237009377 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27242563991 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27242563991 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27253510982 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27253510982 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27253510982 |
| image: lmsysorg/sglang:nightly-dev-cu13-20260603-83bc7766 | ||
| model: deepseek-ai/DeepSeek-V4-Pro | ||
| model-prefix: dsv4 | ||
| runner: gb300-nv |
There was a problem hiding this comment.
Wrong runner for SGLang recipes
High Severity
The new dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k entry uses runner: gb300-nv, while sibling DeepSeek-V4 GB300 dynamo-sglang configs use gb300-cw. launch_gb300-nv.sh never copies staged recipes/sglang/deepseek-v4 into srt-slurm (only glm5 gets that path), and its srtslurm.yaml omits the dsv4-pro alias many new recipes use—so srtctl apply is likely to fail on missing recipes or model preflight.
Reviewed by Cursor Bugbot for commit 47460ef. Configure here.
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27366225297 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27366364441 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27372623348 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=27372623348 |
|
@hshrivastava-droid The GB300 CW are back |
…to existing 8k1k key Move the 11 isl=1024 fixed-seq-len entries out of the standalone dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k key and append them to the existing dsv4-fp4-gb300-dynamo-sglang-mtp key so a single top-level config covers both 8k/1k and 1k/1k scenarios on GB300. Master config image stays at lmsysorg/sglang:nightly-dev-20260527-14f81a67 (unchanged); the 11 new 1k/1k recipes remain pinned to lmsysorg/sglang:v0.5.13.post1-cu130 per PR head.
Restore the 7 disagg-gb300-*.yaml recipes referenced by the non-MTP dsv4-fp4-gb300-dynamo-sglang config to their main-branch container tag. These bumps were unrelated to the -mtp key consolidation in this PR.
Resolve conflicts: - perf-changelog.yaml: take main's version and append the PR 1697 entry (dsv4-fp4-gb300-dynamo-sglang-mtp) at the end - runners/launch_gb300-nv.sh: keep both concurrent dynamo-sglang elif branches (our dsv4 + main's qwen3.5)
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=28613151268 |
Previous tag lmsysorg/sglang:nightly-dev-20260527-14f81a67 expired on Docker Hub (404), causing every job in run 28613151268 (PR #1697) to fail at enroot import. Bump the 8 MTP recipes and the nvidia-master entry to lmsysorg/sglang:nightly-dev-cu13-20260706-8673e85e (verified live). Update the PR's perf-changelog entry to reflect the bump. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
# Conflicts: # perf-changelog.yaml # runners/launch_gb300-nv.sh
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=28822342442 |
1 similar comment
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=28822342442 |
…26) for dsv4+dynamo-sglang
The new 1k/1k STP recipes use benchmark: {type: custom, command: ...},
a schema feature that only exists on NVIDIA/srt-slurm main. Pinning
sa-submission-q2-2026 caused srtctl to reject the recipe with
"Invalid config ... {'benchmark': {'command': ['Unknown field.']}}"
before any benchmark could run (see failing sweep run 28977862941).
Same launcher fix PR #1697 already carries; applying it here so the
dynamo-sglang + dsv4 elif clones NVIDIA/srt-slurm@main.
Pin master image: and all 19 (11 x 1k1k + 8 x 8k1k) recipe model.container values to lmsysorg/sglang:nightly-dev-cu13-20260710-cfc66e05, aligning the whole config on the newer nightly and clearing the prior split between 1k1k (v0.5.13.post1-cu130) and 8k1k (07-06 nightly).
The runners/launch_gb300-nv.sh dynamo-sglang + dsv4 branch clones NVIDIA/srt-slurm and previously ran 'git checkout main', so every sweep picked up whatever srtctl 'main' happened to be at that moment. Between the last green run (2026-06-16) and the current failing head (2026-07-06), main advanced meaningfully, which is a plausible driver of the recent 'decode worker exit 137 / Server did not become healthy' failures on the byte-identical 1k1k recipes. Pin to v1.0.17 so all future sweeps in this PR run against a fixed srtctl SHA and drift can be ruled in or out cleanly.
Resolve the PR #1697 merge conflicts by keeping main's changelog entries and appending the branch entry, while preserving both launcher paths. Also remove recipe-level 3h Slurm limits from the new 1k1k disagg recipes so GitHub Actions' 8h cap governs the sweep. 中文:将 main 同步到 dsv4 GB300 配置分支。通过保留 main 的 changelog 条目并在末尾追加本分支条目来解决 PR #1697 的合并冲突,同时保留两个启动器路径;并移除新增 1k1k 分离式配方中的 3 小时 Slurm 限制,让 GitHub Actions 的 8 小时上限控制扫描运行。
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29129150548 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29129150548 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29129150548 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29540585164 |
2 similar comments
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29540585164 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29540585164 |
|
/reuse-sweep-run |
|
/reuse-sweep-run 29540585164 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — |


Note
Low Risk
Benchmark and CI launcher configuration only; no application runtime or auth paths change, though miswired recipe paths could waste cluster jobs.
Overview
Adds a new
dsv4-fp4-gb300-dynamo-sglang-mtp-1k1kentry innvidia-master.yaml(separate from the existing 8k/1k sibling because the launcher pins one container image per top-level config). It targets DeepSeek-V4-Pro FP4 on GB300 with disaggregated SGLang + MTP, 1k ISL / 1k OSL, and wires 11CONFIG_FILEscenarios: three conc=8192 throughput shapes (dep8 / dep16 / 2p1d-dep16) and eight low-latency splits (dep4 vs tp4 prefill/decode, with 1/2/4/6 decode workers).Stages matching Slurm recipes under
benchmarks/multi_node/srt-slurm-recipes/sglang/deepseek-v4/1k1k/(Dynamo + multi-frontend for high-conc; SGLang frontend for low-lat). Decode uses EAGLE speculative decoding on all new recipes.Also bumps several existing 8k1k GB300 recipe
model.containervalues from a nightly tag tolmsysorg/sglang:v0.5.13.post1-cu130, documents the work inperf-changelog.yaml, and extendsrunners/launch_gb300-nv.shso dynamo-sglang + dsv4 clones NVIDIA/srt-slurmmainand overlays the localdeepseek-v4recipe tree at job time.Reviewed by Cursor Bugbot for commit 90cee64. Bugbot is set up for automated code reviews on this repo. Configure here.
中文说明
在
nvidia-master.yaml中新增dsv4-fp4-gb300-dynamo-sglang-mtp-1k1k配置条目,用于 DeepSeek-V4-Pro FP4 在 GB300 上通过 Dynamo + SGLang + MTP 进行分离式推理基准测试(ISL 1k / OSL 1k)。包含 11 个CONFIG_FILE场景:3 个高吞吐拓扑(conc=8192)和 8 个低延迟拓扑。所有新配方均使用 EAGLE 投机解码。同时将若干已有 8k1k GB300 配方的容器镜像版本升级至lmsysorg/sglang:v0.5.13.post1-cu130,并更新了perf-changelog.yaml和launch_gb300-nv.sh启动脚本。