-
Notifications
You must be signed in to change notification settings - Fork 233
Add GLM-5.2 NVFP4 B300 SGLang single-node agentic benchmarks #2268
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from 1 commit
264f2ae
e46fae9
8933d94
c833882
3643eb2
21d9de2
c21ff08
53fa0bb
0ac5683
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,150 @@ | ||
| #!/usr/bin/env bash | ||
| set -euo pipefail | ||
| set -x | ||
|
|
||
| # Agentic trace replay benchmark for GLM-5.2 NVFP4 on B300 using SGLang. | ||
|
Check warning on line 5 in benchmarks/single_node/agentic/glm5.2_fp4_b300_sglang.sh
|
||
| # | ||
| # Server flags follow the SGLang cookbook B300 NVFP4 single-node recipes | ||
| # (https://docs.sglang.io/cookbook/autoregressive/GLM/GLM-5.2), STP only: | ||
| # the cookbook's EAGLE MTP variants are intentionally not wired up yet. | ||
| # DP_ATTENTION=false -> low-latency arm (TP8, fp8 KV, cutedsl bf16 GEMM) | ||
| # DP_ATTENTION=true -> high-throughput arm (TP8 + DP8 attention-DP) | ||
| # | ||
| # Required env vars: | ||
| # MODEL, TP, CONC, KV_OFFLOADING, TOTAL_CPU_DRAM_GB, RESULT_DIR, DURATION, | ||
| # EP_SIZE, DP_ATTENTION | ||
|
|
||
| source "$(dirname "$0")/../../benchmark_lib.sh" | ||
|
|
||
| check_env_vars MODEL TP CONC KV_OFFLOADING TOTAL_CPU_DRAM_GB RESULT_DIR DURATION EP_SIZE DP_ATTENTION | ||
|
|
||
| if [[ "$KV_OFFLOADING" != "none" ]]; then | ||
| echo "Error: KV_OFFLOADING=$KV_OFFLOADING is not supported by this recipe" >&2 | ||
| exit 1 | ||
| fi | ||
|
|
||
| if [[ -n "${SLURM_JOB_ID:-}" ]]; then | ||
| echo "JOB $SLURM_JOB_ID running on ${SLURMD_NODENAME:-unknown}" | ||
| fi | ||
|
|
||
| # `hf download` creates the target dir if missing and is itself idempotent. | ||
| # When MODEL_PATH is unset (stand-alone runs), fall back to the HF_HUB_CACHE. | ||
| # Either way, MODEL_PATH is what the server is launched with. | ||
| if [[ -n "${MODEL_PATH:-}" ]]; then | ||
| if [[ ! -d "$MODEL_PATH" || -z "$(ls -A "$MODEL_PATH" 2>/dev/null)" ]]; then | ||
| hf download "$MODEL" --local-dir "$MODEL_PATH" | ||
| fi | ||
| else | ||
| hf download "$MODEL" | ||
| export MODEL_PATH="$MODEL" | ||
| fi | ||
| nvidia-smi | ||
|
|
||
| resolve_trace_source | ||
| install_agentic_deps | ||
|
|
||
| SERVER_LOG="$RESULT_DIR/server.log" | ||
| mkdir -p "$RESULT_DIR" | ||
|
|
||
| # With attention-DP, front the DP ranks with sglang-router using consistent | ||
| # hashing on the AIPerf correlation id so multi-turn sessions stay on the DP | ||
| # rank that holds their radix-cache prefix. | ||
| USE_SGLANG_ROUTER=false | ||
| SGLANG_BACKEND_PORT="$PORT" | ||
| ROUTER_LOG="$RESULT_DIR/router.log" | ||
| if [ "$DP_ATTENTION" = "true" ]; then | ||
| USE_SGLANG_ROUTER=true | ||
| export AIPERF_HTTP_X_SMG_ROUTING_KEY_FROM_CORRELATION_ID=true | ||
| SGLANG_BACKEND_PORT=$((PORT + 1)) | ||
| SGLANG_ROUTER_METRICS_PORT=$((PORT + 10000)) | ||
| fi | ||
|
|
||
| PARALLEL_ARGS=(--tp "$TP" --ep-size "$EP_SIZE") | ||
| if [ "$DP_ATTENTION" = "true" ]; then | ||
| PARALLEL_ARGS+=( | ||
| --dp "$TP" | ||
| --enable-dp-attention | ||
| --tokenizer-worker-num "$TP" | ||
| --dist-init-addr "127.0.0.1:$((PORT + 2000))" | ||
| ) | ||
| else | ||
| # Cookbook low-latency levers; the DP-attention cell omits them. | ||
| PARALLEL_ARGS+=( | ||
| --kv-cache-dtype fp8_e4m3 | ||
| --bf16-gemm-backend cutedsl | ||
| --max-prefill-tokens 8192 | ||
| ) | ||
| fi | ||
|
|
||
| # AgentX concurrency counts live session trees, not individual requests. | ||
| # Allow subagent fan-out to exceed CONC without clipping request bursts. | ||
| MAX_RUNNING_REQUESTS=$((2 * CONC)) | ||
| GRAPH_ARGS=() | ||
| if [ "$DP_ATTENTION" != "true" ]; then | ||
| # Cookbook low-latency captures graphs up to its request cap; the | ||
| # DP-attention cell leaves the CUDA-graph batch list at SGLang defaults. | ||
| CUDA_GRAPH_MAX_BS=$MAX_RUNNING_REQUESTS | ||
| [ "$CUDA_GRAPH_MAX_BS" -gt 64 ] && CUDA_GRAPH_MAX_BS=64 | ||
| GRAPH_ARGS=(--cuda-graph-max-bs "$CUDA_GRAPH_MAX_BS") | ||
| fi | ||
|
|
||
| export PYTHONNOUSERSITE=1 | ||
| export TORCH_CUDA_ARCH_LIST=10.0 | ||
| # Agentic warmup dispatches hundreds of large prompts at once; allow up to | ||
| # 15 minutes of TCP progress before AIPerf declares a connection dead. | ||
| export AIPERF_HTTP_TCP_USER_TIMEOUT=900000 | ||
|
|
||
| SGLANG_CMD=( | ||
| python3 -m sglang.launch_server | ||
| --model-path "$MODEL_PATH" | ||
| --served-model-name "$MODEL" | ||
| --host 0.0.0.0 | ||
| --port "$SGLANG_BACKEND_PORT" | ||
| --trust-remote-code | ||
| "${PARALLEL_ARGS[@]}" | ||
| --quantization modelopt_fp4 | ||
| --chunked-prefill-size 8192 | ||
| --mem-fraction-static 0.85 | ||
| --max-running-requests "$MAX_RUNNING_REQUESTS" | ||
| "${GRAPH_ARGS[@]}" | ||
| --watchdog-timeout 1800 | ||
| --enable-metrics | ||
| ) | ||
|
|
||
| printf '%q ' "${SGLANG_CMD[@]}" | tee "$RESULT_DIR/sglang_command.txt" | ||
| printf '\n' | tee -a "$RESULT_DIR/sglang_command.txt" | ||
|
|
||
| echo "Starting SGLang server for B300..." | ||
| "${SGLANG_CMD[@]}" > "$SERVER_LOG" 2>&1 & | ||
| SERVER_PID=$! | ||
| echo "Server PID: $SERVER_PID" | ||
|
|
||
| wait_for_server_ready --port "$SGLANG_BACKEND_PORT" --server-log "$SERVER_LOG" --server-pid "$SERVER_PID" | ||
|
|
||
| if [ "$USE_SGLANG_ROUTER" = "true" ]; then | ||
| echo "Starting SGLang router on port $PORT for $TP DP ranks..." | ||
| python3 -m sglang_router.launch_router \ | ||
| --worker-urls "http://localhost:$SGLANG_BACKEND_PORT" \ | ||
| --policy consistent_hashing \ | ||
| --request-id-headers x-correlation-id \ | ||
| --dp-aware \ | ||
| --host 0.0.0.0 \ | ||
| --port "$PORT" \ | ||
| --prometheus-host 127.0.0.1 \ | ||
| --prometheus-port "$SGLANG_ROUTER_METRICS_PORT" \ | ||
| --connect-timeout-secs 900 \ | ||
| --request-timeout-secs 14400 \ | ||
| --disable-health-check \ | ||
| --disable-retries > "$ROUTER_LOG" 2>&1 & | ||
| ROUTER_PID=$! | ||
| echo "Router PID: $ROUTER_PID" | ||
| wait_for_server_ready --port "$PORT" --server-log "$ROUTER_LOG" --server-pid "$ROUTER_PID" | ||
| fi | ||
|
|
||
| if [ "${EVAL_ONLY}" = "true" ]; then | ||
| run_eval --port "$PORT" | ||
| else | ||
| build_replay_cmd "$RESULT_DIR" | ||
| REPLAY_CMD+=" --server-metrics http://localhost:$SGLANG_BACKEND_PORT/metrics" | ||
| run_agentic_replay_and_write_outputs "$RESULT_DIR" | ||
| fi | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -9891,3 +9891,25 @@ | |
| tp: 16 | ||
| ep: 16 | ||
| dp-attn: true | ||
|
|
||
| # GLM-5.2 B300 NVFP4 AgentX frontier from the SGLang cookbook B300 NVFP4 | ||
| # single-node recipes (https://docs.sglang.io/cookbook/autoregressive/GLM/GLM-5.2), | ||
| # STP only (the cookbook's EAGLE MTP variants are deferred). The TP8 arm is the | ||
| # cookbook low-latency recipe (fp8 KV, cutedsl bf16 GEMM) and covers the | ||
| # interactivity end; the TP8/DP8 attention-DP arm is the cookbook | ||
| # high-throughput recipe behind sglang-router consistent hashing for session | ||
| # affinity. Conc lists are disjoint between arms so exp-names stay unique. | ||
| glm5.2-fp4-b300-sglang-agentic: | ||
| image: lmsysorg/sglang:v0.5.15.post1-cu130 | ||
| model: nvidia/GLM-5.2-NVFP4 | ||
| model-prefix: glm5.2 | ||
| runner: cluster:b300-nv | ||
| precision: fp4 | ||
| framework: sglang | ||
| multinode: false | ||
| scenarios: | ||
| agentic-coding: | ||
| - dram-utilization: 0.80 | ||
| search-space: | ||
| - { tp: 8, kv-offloading: none, conc-list: [1, 2, 4, 8, 16, 32] } | ||
| - { tp: 8, dp-attn: true, kv-offloading: none, conc-list: [48, 64, 96, 128, 192, 256, 512], router: { name: sglang-router, version: "0.3.2" } } | ||
|
Check failure on line 9915 in configs/nvidia-master.yaml
|
||
|
Comment on lines
+7864
to
+7877
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. 🔴 The high-throughput arm ( Extended reasoning...The bug: The high-throughput search-space entry for - { tp: 8, dp-attn: true, kv-offloading: none, conc-list: [48, 64, 96, 128, 192, 256, 512], router: { name: sglang-router, version: "0.3.2" } }It sets Fields.EP.value: ep if ep is not None else 1,Since the config never supplies Why this is wrong: GLM-5.2 is an MoE model ( Why nothing catches this today: There's no validation in Comparison to every other recipe in the file: Grepping Step-by-step proof:
Fix: add - { tp: 8, ep: 8, dp-attn: true, kv-offloading: none, conc-list: [48, 64, 96, 128, 192, 256, 512], router: { name: sglang-router, version: "0.3.2" } } |
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
🟡 This PR's title and description are English-only, violating the AGENTS.md rule that all PR titles/descriptions must be bilingual (title format ' / <中文标题>' plus a mirrored Chinese section in the body).
Extended reasoning...
AGENTS.md line 7 states as an explicit, mandatory rule: "PR and GitHub-issue titles & descriptions must be bilingual — include a Simplified Chinese version in addition to English. Title format:
<English title> / <中文标题>. In the PR/issue body, follow the English content with its Chinese translation (e.g. a## 中文说明section mirroring the summary...)". This applies to every PR, with no carve-out for benchmark-recipe PRs.This PR's title is 'Add GLM-5.2 NVFP4 B300 SGLang single-node agentic benchmarks' — no '/ <中文标题>' suffix — and the description's Summary, Changes, and Validation sections are entirely in English with no mirrored '## 中文说明' section or equivalent.
This is not a matter of subjective style: the repo's own commit log shows a sibling recipe PR following the rule correctly. Commit d85fa13 (PR #2182) has the bilingual title 'Add MiniMax M3 8k/1k Dynamo vLLM B300 EAGLE recipes / 新增 MiniMax M3 8k/1k Dynamo vLLM B300 EAGLE 配方', proving other contributors in this exact benchmark-recipe workflow are expected to (and do) satisfy the rule. This PR is the outlier.
Step-by-step proof:
Impact: none on benchmark correctness or CI — this is a process/documentation compliance gap, not a code defect. Fix is simple: the author (or whoever finalizes the merge) should append '/ <一句中文标题>' to the PR title and add a '## 中文说明' section mirroring the Summary/Changes/Validation content before merge.