Skip to content

Commit c84cf75

Browse files
User Nameclaude
andcommitted
397B README: forward-link the FP8 redo + microbench-index taxonomy note
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
1 parent 5cc4ba2 commit c84cf75

1 file changed

Lines changed: 3 additions & 1 deletion

File tree

  • hardware-tests/qwen3.5-397b-vs-step3.7-flash-2026-05-29

hardware-tests/qwen3.5-397b-vs-step3.7-flash-2026-05-29/README.md

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,8 @@ Step-3.7-Flash-NVFP4 entry on the same box.
66

77
**N=10** — ten replicates per cell, both arms (240 cells, all `done_signal`; phase-1 graded with the fixed `phase1_grade.py`).
88

9+
> **Where this lives:** this is a **model-behavior** study (the 12-family agentic microbench), physically filed under `hardware-tests/` only because it needed the dual-Blackwell rig. For every 12-family microbench across both trees, see [`../../MICROBENCH-INDEX.md`](../../MICROBENCH-INDEX.md).
10+
911
## TL;DR
1012

1113
*This entry is methodological, not "which model won." The two results that survive scrutiny lead; the "scale ties" observation is real but the most caveated, so it's demoted.*
@@ -15,7 +17,7 @@ Step-3.7-Flash-NVFP4 entry on the same box.
1517
- **③ Aggregate ties ~7–8/12 across 397B / Flash / 27B-Q4 / Coder-Q4 — but read as *suggestive*.** Two confounds keep this from being a scaling law: **cross-quant** (397B at Q3 vs ~11B-active at FP4 — not a clean scale axis) and **N-asymmetry** (only 397B is N=10; comparators are N=1, which this very entry proves misreads cells).
1618
- **Failure temperament tracks lineage, not size:** 397B + 27B *stall* (never over-generate); Coder-Next + Flash *run away*. Zero max_tokens runaways across all 240 397B cells.
1719
- ⚠️ 27B/Coder **phase-1 reference cells are quarantined** pending [issue #29] (same grader bug this entry fixed); their p2/p3 cells are unaffected and used in the cross-model comparison.
18-
- **Cross-model uses clean Q4/AWQ refs** for 27B/Coder; fresh Q8/FP8 runs excluded as serving failures (documented, not faked).
20+
- **Cross-model uses clean Q4/AWQ refs** for 27B/Coder; the *first* fresh Q8/FP8 attempts were excluded as serving failures (documented, not faked). **Update (2026-05-31):** a clean **FP8** redo of 27B since succeeded — full entry at [`../qwen3.6-27b-fp8-microbench-2026-05-31/`](../qwen3.6-27b-fp8-microbench-2026-05-31/); the failure was **Q8-serving-specific**, not 27B-on-this-rig.
1921
- **GPU power:** combined both-GPU draw never within 5% of the 1200W cap (median 670W, max 985W=82%); GPU0 leads GPU1 — pipeline alternation. The pair never hits full power together.
2022
- The substance is qualitative — **read [QUALITATIVE.md](QUALITATIVE.md).**
2123

0 commit comments

Comments
 (0)