/overview data gaps aren't a timeout problem — normal benchmarks run ~15–80 min against a 480–500 min cap. The gap is that hardware isn't run in the same batch, so cross-hardware numbers aren't strictly comparable.
Not the fix: raising CI runtime/timeout.
Approach — a calibration workflow: pin one snapshot (same commit / image / model / config / workload), run it across the exec-relevant hardware, and publish the result as a single dated snapshot. Pilot small first; expand only if the pilot proves comparability. Pilot scope TBD in thread.
Next phase for SemiAnalysisAI/InferenceX-app#611 — not a blocker for the current /overview PR.
/overview 的数据缺口不是超时问题——常规 benchmark 只跑约 15–80 分钟,而上限是 480–500 分钟。真正的缺口在于不同硬件没有在同一批次运行,因此跨硬件的数字并不严格可比。
不是解决方案:提高 CI 运行时长/超时上限。
方案——校准 workflow:固定一个 snapshot(相同的 commit/镜像/模型/配置/workload),在面向高层的相关硬件上运行,并将结果作为一个带日期的 snapshot 发布。先做小规模 pilot;仅当 pilot 证明可比性后再扩展。Pilot 范围在 issue 讨论中再定。
SemiAnalysisAI/InferenceX-app#611 的下一阶段——不阻塞当前 /overview PR。
/overviewdata gaps aren't a timeout problem — normal benchmarks run ~15–80 min against a 480–500 min cap. The gap is that hardware isn't run in the same batch, so cross-hardware numbers aren't strictly comparable.Not the fix: raising CI runtime/timeout.
Approach — a calibration workflow: pin one snapshot (same commit / image / model / config / workload), run it across the exec-relevant hardware, and publish the result as a single dated snapshot. Pilot small first; expand only if the pilot proves comparability. Pilot scope TBD in thread.
Next phase for SemiAnalysisAI/InferenceX-app#611 — not a blocker for the current
/overviewPR./overview的数据缺口不是超时问题——常规 benchmark 只跑约 15–80 分钟,而上限是 480–500 分钟。真正的缺口在于不同硬件没有在同一批次运行,因此跨硬件的数字并不严格可比。不是解决方案:提高 CI 运行时长/超时上限。
方案——校准 workflow:固定一个 snapshot(相同的 commit/镜像/模型/配置/workload),在面向高层的相关硬件上运行,并将结果作为一个带日期的 snapshot 发布。先做小规模 pilot;仅当 pilot 证明可比性后再扩展。Pilot 范围在 issue 讨论中再定。
SemiAnalysisAI/InferenceX-app#611 的下一阶段——不阻塞当前
/overviewPR。