Commit ab024f1
committed
bench(gb10): refresh 1B matrix — path_b now runs on CUDA (was blocked)
After the path_b CUDA fix (a319599), all bf16+fp8 path_b cells run on gb10:
bf16 path_b tok/s muon=92 adamw=156 lion=168 (faster than path_c at bs=1 —
CUDA-eager reference is leaner than fused for tiny batch), losses 11.3->6.27 PASS.
36 ok (bf16+fp8 x path_b/path_c/path_c_chunked), 18 nvfp4 blocked (fail-loud,
no nvfp4 training kernels). Full 54-cell matrix, memory-guarded run (<70G cap).1 parent 82aa315 commit ab024f1
2 files changed
Lines changed: 99 additions & 99 deletions
0 commit comments