You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/evidence/experiment-master-log.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -13,7 +13,7 @@
13
13
| Item | Verdict | Evidence |
14
14
| --- | --- | --- |
15
15
|**Phase G H1 historical baseline**|**Quarantined.** Local seed42/43/45 targets trained on rows labeled as nonmembers, and the old H1 scorer reported resubstitution AUC. All old AUCs, continuation directions, N=512 clusters, knockout/fine-grid interpretations, and mechanism claims are diagnostic-only. Never resume these checkpoints into corrected runs. |[frozen-claim-matrix.md](../paper1/frozen-claim-matrix.md)|
16
-
|**Paper 1 corrected-evidence rebuild**|Active protocol, no result yet. Freeze member-only training, predeclared seeds, disjoint calibration/evaluation rows, common noise, H1 held-out scoring, PIA positive control, and STOP/REPLICATE/MATURE rules before viewing any corrected outcome. |[paper1-corrected-evidence-runbook-2026-07-11.md](../start-here/paper1-corrected-evidence-runbook-2026-07-11.md)|
16
+
|**Paper 1 corrected-evidence rebuild**|Batch-32 training and exact-resume resource preflight passed. The first formal matrix was then stopped outcome-blind at step 22,000 on its first target after a post-training audit found that canonical PIA could not load corrected `ema` checkpoints and lacked the frozen complete-roster statistics. No membership outcome was generated or viewed. The partial target is permanently excluded; corrected EMA loading, analysis provenance, canonical roster ordering, PIA confirmatory statistics, and 1,000-draw H1/PIA paired bootstrap are being refrozen before a fresh restart. |[paper1-corrected-training-preflight-2026-07-12.md](paper1-corrected-training-preflight-2026-07-12.md)|
17
17
| DualMD / DistillMD defense artifact gate | OpenReview DDMD supplement code and DDPM split-index files are public, but the embedded GitHub origin is not public and no checkpoint-bound defended/undefended scores, ROC arrays, metric JSON, generated response packet, or ready verifier are committed; no download, GPU release, or admitted defense row |[dual-md-distillmd-defense-artifact-gate-2026-05-15.md](dualmd-distillmd-defense-artifact-gate-2026-05-15.md)|
18
18
| DIFFENCE classifier-defense artifact gate | official code, configs, and split-index files are public, but the protected target is an image classifier, diffusion is a pre-inference defense component, and no checkpoint-bound defended/undefended logits, score rows, ROC arrays, metric JSON, or ready verifier are committed; no download, GPU release, or admitted defense row |[diffence-classifier-defense-artifact-gate-2026-05-15.md](diffence-classifier-defense-artifact-gate-2026-05-15.md)|
19
19
| MIAHOLD / HOLD++ higher-order Langevin defense gate | official defense code, audio split filelists, CIFAR HOLD config, and PIA-style attack code are public, but checkpoint-bound target artifacts, reusable member/nonmember scores, ROC arrays, metric JSON, generated responses, and a ready verifier are missing; no download, GPU release, or admitted defense row |[miahold-higher-order-langevin-artifact-gate-2026-05-15.md](miahold-higher-order-langevin-artifact-gate-2026-05-15.md)|
-**Status**: superseded by outcome-blind evaluation-contract repair
5
+
-**Verdict**: Batch size 32 passed the training and exact-resume resource gate, but the first formal matrix was stopped before outcomes when the canonical PIA post-training path was found incomplete.
6
+
7
+
## Scope
8
+
9
+
This preflight tested only the corrected member-only training contract and its
10
+
artifact chain. It did not inspect H1, PIA, AUC, or any other membership outcome.
11
+
The run used the first predeclared seed, 25,000 fixed member rows, the frozen
12
+
split SHA256, and corrected-only output labels.
13
+
14
+
## Resource correction before outcomes
15
+
16
+
The initial batch-64 attempt was stopped before completion. Observed GPU memory
17
+
reached about 7.9 GiB on an 8 GiB device, leaving only tens of MiB free, and
18
+
2,000 steps had not completed after about 22 minutes. This failed the frozen
19
+
6.0k steps/hour and VRAM-headroom gates.
20
+
21
+
Before any corrected metric was viewed, the canonical training config was
22
+
changed once to batch size 32. No model, optimizer, schedule, precision, seed,
23
+
or step horizon was changed. The protocol was then rebuilt and hash-sealed.
24
+
25
+
## Passing run
26
+
27
+
The replacement run used three segments to exercise resume:
28
+
29
+
| Segment | Wall time | Approx. throughput |
30
+
| --- | ---: | ---: |
31
+
| 0 -> 200 | 57 s | 12.6k steps/hour |
32
+
| 200 -> 400 | 67 s | 10.7k steps/hour, including resume overhead |
33
+
| 400 -> 2,000 | 459 s | 12.5k steps/hour |
34
+
| Active total | 583 s | 12.3k steps/hour |
35
+
36
+
Observed runtime bounds:
37
+
38
+
- fresh-run GPU allocation: about 5.50 GiB;
39
+
- resume peak GPU allocation: about 6.74 GiB;
40
+
- minimum observed free GPU memory: about 1.21 GiB;
41
+
- maximum observed temperature: 71 C;
42
+
- maximum observed power draw: about 103 W;
43
+
- no OOM or observed thermal throttling.
44
+
45
+
## Artifact checks
46
+
47
+
The run produced checkpoints at steps 200, 400, and 2,000. For every checkpoint:
48
+
49
+
- the on-disk SHA256 matches its manifest receipt;
50
+
- checkpoint metadata records the same protocol hash, split SHA256, code commit,
51
+
seed, step, and batch-32 training-config hash;
52
+
- the manifest retains all three receipts after resume;
53
+
-`weights_only=True` loading succeeds on CPU;
54
+
- log terminal events report the expected final step without interruption.
55
+
56
+
Generated checkpoints and logs remain outside Git under the configured download
57
+
and training-output roots.
58
+
59
+
## Decision
60
+
61
+
The corrected training implementation passes Stage 0's training, resource,
62
+
resume, and receipt gates. This does not release the four 100k targets. The next
63
+
required gate is a complete evaluation benchmark on the 2,000-step checkpoint:
64
+
65
+
1. extract H1 activations only for the locked calibration/evaluation rows;
66
+
2. derive common noise from the sealed row/timestep/draw contract;
67
+
3. emit strict row-bound H1 and PIA score packets;
68
+
4. time the full H1 + PIA evaluation and project the worst non-STOP branch with
69
+
a 10% failure buffer.
70
+
71
+
No long training begins until that projection fits the available GPU window.
72
+
73
+
## Outcome-blind formal-matrix abort
74
+
75
+
The full H1 and canonical PIA timing projection subsequently fit the available
76
+
window. A formal Stage 1 launch therefore began on the frozen target matrix.
77
+
During an independent audit of the commands that would consume the completed
78
+
100k checkpoints, two remaining P0 gaps were found before any membership score
79
+
or metric was generated:
80
+
81
+
1. the formal canonical PIA exporter accepted legacy `ema_model`/`net_model`
82
+
checkpoint keys but not the corrected trainer's `ema` key;
83
+
2. the PIA path did not yet expose the predeclared complete-roster paired
84
+
bootstrap, global AUC-range test, and Holm family required for the paper
85
+
decision.
86
+
87
+
The first target was stopped at step 22,000. Its final saved checkpoint receipt
88
+
matched its on-disk SHA256; the other three targets had not started. This partial
89
+
target is an infrastructure-abort artifact and MUST NOT be resumed or scored.
90
+
91
+
Because no corrected outcome had been viewed, the protocol was reopened rather
92
+
than patched after results. The repair adds corrected EMA loading, explicit
93
+
analysis repository-state provenance, canonical target ordering, PIA
94
+
confirmatory statistics, and 1,000 paired bootstrap draws for both H1 and PIA.
95
+
The increase from 200 to 1,000 is required because 200 plus-one draws make Holm
96
+
rejection mathematically unreachable for the 28 pairs in an eight-target set.
97
+
98
+
Long training remains blocked until the repaired code is committed, a new
99
+
protocol hash is sealed, and the necessary dry-run/resource preflight is
100
+
repeated with fresh output labels. The aborted target and its protocol are never
0 commit comments