Skip to content

Commit 7af9080

Browse files
author
gHashTag
committed
repro/numerics: measured L4/L5 GF16 vs bf16/fp16 accuracy oracle
Add repro/numerics/{nmse_gf16.py,nmse_manifest.json,README.md} and fill the L4/L5 TBD cells in docs/NUMERICS_VALIDATION.md with measured, reproducible round-trip NMSE/ULP numbers for GF16 (E6M9) vs bfloat16 and float16, per docs/GF16_BFLOAT16_NMSE_PROTOCOL.md. Host-measured, unsealed => informational, not a silicon certifying claim. D_PHI is an identity sanity check. Closes #1072
1 parent 80c7b48 commit 7af9080

5 files changed

Lines changed: 507 additions & 15 deletions

File tree

docs/NOW.md

Lines changed: 7 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,6 +1,12 @@
11
# NOW -- Trinity t27 sync
22

3-
Last updated: 2026-06-13
3+
Last updated: 2026-06-14
4+
5+
## repro-numerics-l4-l5-oracle -- measured GF16 vs bf16/fp16 round-trip NMSE/ULP (Closes #1072)
6+
7+
- **WHERE** (`repro/numerics/` + `docs/NUMERICS_VALIDATION.md`): adds `repro/numerics/nmse_gf16.py`, `repro/numerics/nmse_manifest.json`, and `repro/numerics/README.md`. The script round-trips real -> format -> real through the GF16 codec SSOT (`conformance/gf16_ref.py`, E6M9) and through `ml_dtypes.bfloat16` and `numpy.float16`, computing NMSE, NMSE ratios, max-abs-err, a ULP-like (kappa-approx) metric, and an overflow/saturation rate over the protocol's reference distributions (`D_NORM`, `D_LOG`, `D_RELU`, `D_PHI`, `D_DEEP`) plus a documented wide-range extension `D_WIDE` (log2|x| ~ U(-28,28)). It enforces the L5 identity witness (phi^2=phi+1, phi^2+phi^-2=3) before any number and aborts non-zero on witness failure or negative NMSE. Replaces the `TBD` cells in `docs/NUMERICS_VALIDATION.md` sections 4 (ladder L4/L5), 5 (differential oracle table), 6 (IEEE/bf16 baseline) with measured values; updates section 11 reproduction and the closing line.
8+
- **Why**: `docs/NUMERICS_VALIDATION.md` named the missing differential/comparative oracle as the project's predictable-skepticism gap (L4/L5 = TBD/P1). This closes it at host level with reproducible numbers (seed 2718281, 2,000,000 samples/distribution): GF16 NMSE is ~16x lower than bf16 (ratio ~0.063, 9 vs 7 mantissa bits) and ~4x higher than fp16 (10 vs 9 mantissa bits) near 1.0, while on D_WIDE fp16 overflows 21.4% of samples and GF16 (max ~4.29e9) and bf16 lose none -- the honest exponent/mantissa tradeoff predicted by protocol section 3.3 and matching the E6M9 motivation of IBM DLFloat (DOI 10.1109/ARITH.2019.00023) and Popescu et al. (arXiv:2103.15940). R5-HONEST: host-measured, unsealed => informational, not a silicon certifying claim; D_PHI is an identity sanity check, not a superiority claim. ASCII-only. L6 SSOT untouched. Context preprint: arXiv:2606.05017. Closes #1072.
9+
- **Anchor**: phi^2 + phi^-2 = 3
410

511
## ops-unify-email -- consolidate contact emails to admin@t27.ai (Closes #1067)
612

docs/NUMERICS_VALIDATION.md

Lines changed: 19 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -44,35 +44,39 @@ Until filled, treat numeric behavior as **implementation-defined** outside confo
4444
| L1 | **Exhaustive** encode/decode + op table | GF4 (and GF8 if feasible) | TBD |
4545
| L2 | **Conformance JSON** — existing `conformance/gf*_vectors.json` | GF4–GF32 as covered | partial |
4646
| L3 | **Property-based / randomized** boundaries | GF16+ | TBD |
47-
| L4 | **Differential** vs reference (Python `decimal`, or MPFR) | GF16 primary | TBD — P1 |
48-
| L5 | **Comparative** vs IEEE fp16 / fp32 / bfloat16 on same corpus | GF16 vs fp16/bf16 | TBD |
47+
| L4 | **Differential** vs reference (round-trip oracle) | GF16 primary | **measured (host, unsealed)**`repro/numerics/` |
48+
| L5 | **Comparative** vs IEEE fp16 / bfloat16 on same corpus | GF16 vs fp16/bf16 | **measured (host, unsealed)**`repro/numerics/nmse_manifest.json` |
4949
| L6 | **Optional** posit reference (where tooling exists) | TBD | TBD |
5050

5151
---
5252

5353
## 5. Differential oracle — skeleton results table
5454

55-
*Replace `TBD` with versioned runs; one row per (format, operation, corpus slice).*
55+
*Measured runs (host, unsealed). Reference oracle = f64 round-trip `real -> format -> real`. Seed 2718281, 2,000,000 samples/distribution. Reproduce: `python repro/numerics/nmse_gf16.py`.*
5656

5757
| Run ID | Format | Operation | Corpus | Reference oracle | Max abs err | ULP-like metric | Pass? | Artifact |
5858
|--------|--------|-----------|--------|------------------|-------------|-----------------|-------|----------|
59-
| TBD | GF16 | add | conformance subset | Python `decimal` | TBD | TBD | TBD | `repro/numerics/` (future) |
60-
| TBD | GF16 | mul | || TBD | TBD | TBD | |
61-
| TBD | GF32 | add | || TBD | TBD | TBD | |
59+
| nmse-2718281-D_NORM | GF16 | round-trip | D_NORM (N(0,1)) | f64 | 3.90e-03 | 3.53e-04 | yes | `repro/numerics/nmse_manifest.json` |
60+
| nmse-2718281-D_LOG | GF16 | round-trip | D_LOG (log2\|x\|~U(-10,10)) | f64 | (see manifest) | (see manifest) | yes | `repro/numerics/nmse_manifest.json` |
61+
| nmse-2718281-D_WIDE | GF16 | round-trip | D_WIDE (log2\|x\|~U(-28,28)) | f64 | (see manifest) | (see manifest) | yes | `repro/numerics/nmse_manifest.json` |
6262

63-
**Falsification:** any cell exceeds stated envelope once §2 is normative → **fail CI** or **downgrade claim** in `RESEARCH_CLAIMS.md`.
63+
**Falsification:** any cell exceeds stated envelope once §2 is normative → **fail CI** or **downgrade claim** in `RESEARCH_CLAIMS.md`. The runner already aborts non-zero if the L5 identity witness fails or any NMSE < 0.
6464

6565
---
6666

6767
## 6. IEEE / bfloat16 baseline — skeleton comparison
6868

6969
Same inputs as §5 where bit patterns map sensibly; document **non-comparable** cases explicitly.
7070

71-
| Metric | GF16 | IEEE fp16 | bfloat16 | IEEE fp32 | Notes |
72-
|--------|------|-----------|----------|-----------|-------|
73-
| Dynamic range (stated) | TBD | TBD | TBD | TBD | From spec / measured |
74-
| MSE on N(0,1) sample | TBD | TBD | TBD | TBD | Trinity Phase-1 style table may be ported |
75-
| Add latency (soft impl) | TBD | TBD || TBD | Host-only; not FPGA |
71+
*Measured, host, unsealed (`repro/numerics/nmse_manifest.json`). Non-comparable cases noted.*
72+
73+
| Metric | GF16 | IEEE fp16 | bfloat16 | Notes |
74+
|--------|------|-----------|----------|-------|
75+
| Mantissa / exponent bits | 9 / 6 | 10 / 5 | 7 / 8 | bit split |
76+
| Max finite magnitude | ~4.29e9 | ~6.55e4 | ~3.39e38 | dynamic range |
77+
| NMSE on N(0,1) (D_NORM) | 1.73e-07 | 4.30e-08 | 2.76e-06 | GF16 ~16x better than bf16; fp16 ~4x better than GF16 |
78+
| Overflow rate on D_WIDE (log2\|x\|~U(-28,28)) | 0.0000 | 0.2144 | 0.0000 | fp16 saturates; GF16/bf16 do not |
79+
| Add latency (soft impl) | n/a | n/a | n/a | host-only round-trip study; latency out of scope (see protocol §1) |
7680

7781
---
7882

@@ -110,8 +114,9 @@ Constant comparisons (if any) must cite **year and revision** and uncertainty; d
110114
## 11. Reproduction
111115

112116
- **Smoke:** `make -C repro repro-numerics` (JSON validity).
113-
- **Future:** `make repro-numerics-diff` (pinned Python + lockfile) — add in `repro/Makefile` when L4 exists.
117+
- **L4/L5 oracle (measured):** `python repro/numerics/nmse_gf16.py` — round-trip NMSE/ULP of GF16 vs bf16/fp16 over the protocol distributions; writes `repro/numerics/nmse_manifest.json`. Host, unsealed.
118+
- **Future:** pinned-toolchain sealed run (`make repro-numerics-diff` with lockfile) for a *certifying* manifest.
114119

115120
---
116121

117-
*Without differential oracles, GoldenFloat will face predictable skepticism — this file is the contract to close that gap.*
122+
*The L4/L5 differential/comparative oracle is now MEASURED (host, unsealed) in `repro/numerics/` — the predictable-skepticism gap is closed at the host level; the remaining step is a sealed-toolchain certifying run.*

repro/numerics/README.md

Lines changed: 77 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,77 @@
1+
# repro/numerics -- L4/L5 differential accuracy oracle for GF16
2+
3+
This directory closes the L4 (differential) and L5 (comparative) rungs of the
4+
numeric validation ladder for GoldenFloat **GF16** (E6M9), per
5+
`docs/GF16_BFLOAT16_NMSE_PROTOCOL.md` and `docs/NUMERICS_VALIDATION.md`.
6+
7+
It produces **measured** round-trip accuracy numbers for GF16 versus
8+
**bfloat16** and **IEEE float16** -- the two 16-bit alternatives of the same
9+
memory footprint -- over the protocol's reference distributions.
10+
11+
> **R5-HONEST.** These are **host-measured, unsealed** numbers. Per protocol
12+
> section 8 they are **informational, not a silicon certifying claim**. No
13+
> result here asserts a measured silicon NMSE for any product. `D_PHI` is an
14+
> identity-anchored sanity check (L5), **not** a superiority claim.
15+
16+
## How to reproduce
17+
18+
```
19+
python repro/numerics/nmse_gf16.py # 2,000,000 samples/distribution
20+
python repro/numerics/nmse_gf16.py --samples 10000000 --seed 2718281
21+
```
22+
23+
- GF16 codec under test: `conformance/gf16_ref.py` (BIAS=31, EXP_BITS=6, MANT_BITS=9).
24+
- bf16: `ml_dtypes.bfloat16`; fp16: `numpy.float16`.
25+
- The run aborts non-zero if the L5 identity witness
26+
(`|phi^2-(phi+1)|<1e-15`, `|phi^2+phi^-2-3|<1e-15`) fails or any NMSE < 0.
27+
- Output manifest: `nmse_manifest.json` (fields per protocol section 6).
28+
29+
## Measured results (seed 2718281, 2,000,000 samples/distribution, unsealed host)
30+
31+
`NMSE(F) = E[(x - Q_F(x))^2] / E[x^2]`. Headline = `NMSE_GF16 / NMSE_BF16`
32+
(< 1 means GF16 is closer to the reference).
33+
34+
| Distribution | NMSE GF16 | NMSE BF16 | NMSE FP16 | GF16/BF16 | GF16/FP16 |
35+
|--------------|-----------|-----------|-----------|-----------|-----------|
36+
| D_NORM | 1.73e-07 | 2.76e-06 | 4.30e-08 | **0.063** | 4.02 |
37+
| D_LOG | 1.48e-07 | 2.35e-06 | 3.68e-08 | **0.063** | 4.02 |
38+
| D_RELU | 1.73e-07 | 2.76e-06 | 4.32e-08 | **0.063** | 4.01 |
39+
| D_PHI | 1.78e-07 | 2.85e-06 | 4.45e-08 | **0.063** | 4.00 |
40+
| D_DEEP | 1.48e-07 | 2.37e-06 | 3.65e-08 | **0.062** | 4.04 |
41+
| D_WIDE | 1.47e-07 | 2.36e-06 | 3.67e-08 | **0.063** | 4.02 |
42+
43+
Saturation / overflow rate (fraction of non-zero samples beyond a format's
44+
finite max: GF16 ~4.29e9, BF16 ~3.39e38, FP16 ~6.55e4):
45+
46+
| Distribution | GF16 | BF16 | FP16 |
47+
|--------------|------|------|------|
48+
| D_WIDE (log2|x| ~ U(-28,28)) | 0.0000 | 0.0000 | **0.2144** |
49+
50+
(all other distributions: 0.0000 for every format)
51+
52+
## Honest interpretation (facts, not claims)
53+
54+
1. **GF16 vs bf16 (mantissa).** GF16 has 9 mantissa bits, bf16 has 7. Across
55+
every distribution GF16's round-trip NMSE is ~16x lower (ratio ~0.063).
56+
This is the expected consequence of the bit split, not a surprise.
57+
2. **GF16 vs fp16 (mantissa).** fp16 has 10 mantissa bits, GF16 has 9, so fp16
58+
is ~4x more accurate near 1.0 (ratio ~4.0). GF16 does **not** beat fp16 on
59+
near-1.0 precision -- stated plainly.
60+
3. **Dynamic range (exponent).** This is where the tradeoff flips. fp16's
61+
exponent saturates at ~65504, so on a wide-range distribution **21.4% of
62+
fp16 samples overflow**, while GF16 (max ~4.29e9, 6-bit exponent) and bf16
63+
(8-bit exponent) lose none. GF16's wider range is the price fp16 pays for
64+
its extra mantissa bit.
65+
4. **Net.** GF16 sits between bf16 and fp16: more precise than bf16, wider
66+
range than fp16. This matches the E6M9 motivation of IBM DLFloat
67+
(ARITH 2019, DOI 10.1109/ARITH.2019.00023) and Popescu et al.
68+
(arXiv:2103.15940), which independently pick a 1/6/9 split.
69+
70+
## Cross-links
71+
72+
- Protocol: `docs/GF16_BFLOAT16_NMSE_PROTOCOL.md`
73+
- Validation ladder: `docs/NUMERICS_VALIDATION.md` (this fills L4/L5)
74+
- Codec SSOT: `conformance/gf16_ref.py`, `conformance/FORMAT-SPEC-001.json`
75+
- Preprint context: arXiv:2606.05017
76+
77+
phi^2 + phi^-2 = 3 | TRINITY

0 commit comments

Comments
 (0)