Commit 05c9e4a
Prometheus: per-row LayerNorm + broadcast-aware tape ops + multi-token attention
The plumbing needed to make multi-token transformer training real.
Rust additions (omnimcode-core/src/interpreter.rs):
(1) tape_layernorm(x, gamma, beta, eps?) — fused per-row LayerNorm
Forward: normalize each row to zero mean / unit variance, scale
by gamma, add beta. Backward: full LayerNorm gradient (dx with
proper centered/scaled terms, dgamma, dbeta).
Composing this from primitives needed broadcast sub/div that
weren't on the tape; fused op is cleaner + faster.
(2) tape_row_mean(x) / tape_row_sum(x) — per-row reductions
[rows, cols] → [rows, 1] with element-wise backward. Building
blocks for any per-row scaling.
(3) tape_add / tape_sub now support row + col vector broadcast
[N, C] + [1, C] (Linear's bias add)
[N, C] + [N, 1] (per-row scaling)
Forward picks the bigger shape; backward reduces upstream gradient
back to the smaller operand's shape via new reduce_to_shape helper.
This fixed a latent bias-gradient bug in the earlier transformer
demo — its 11.3x loss reduction came partly from over-broadcasting
bias grads. With correct broadcast reduction, the same demo gets
4.15x (still real, just honest).
Prometheus additions (examples/lib/prometheus.omc):
(4) prom_layernorm_forward upgraded to use tape_layernorm fused op
instead of the composed (mean, sub, exp(-0.5*log(var+eps)), ...)
path. Cleaner, works on multi-token inputs.
(5) prom_embedding_batch(layer, token_ids[]) — multi-token lookup
via [N, vocab] one-hot @ table. Differentiable into the table.
(6) prom_cross_entropy_batch(logits, targets, vocab) — sum of per-
position -log(softmax) for batched LM training.
A/B demo (examples/prometheus_attention_ab.omc):
Multi-token transformer (8-token windows), seq_len=8, d_model=16,
ff=32, AdamW, cross-entropy. Two arms:
A: alpha=0 (vanilla softmax attention)
B: alpha=0.5 (geodesic-bias attention)
3 seeds × 250 steps each. Tests whether the PyTorch geodesic
win replicates in Prometheus.
The first multi-token training run worked end-to-end through the
new plumbing. Single-seed result before extending to 3 seeds:
vanilla=3.104 geodesic=3.119 delta=+0.46%
A genuine fail-forward. Could be: single seed noise, alpha not
tuned, model too small, training too short. The 3-seed run is
in flight; result will land in the next commit.
What matters infrastructure-wise: multi-token attention works in
pure OMC now. The geodesic primitive is wired correctly (numerically
identical to PyTorch). Whether it HELPS at this scale is an empirical
question we can keep iterating on without re-shipping plumbing.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>1 parent ea3a1e0 commit 05c9e4a
3 files changed
Lines changed: 565 additions & 32 deletions
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
892 | 892 | | |
893 | 893 | | |
894 | 894 | | |
| 895 | + | |
| 896 | + | |
| 897 | + | |
| 898 | + | |
| 899 | + | |
| 900 | + | |
| 901 | + | |
| 902 | + | |
| 903 | + | |
| 904 | + | |
| 905 | + | |
| 906 | + | |
| 907 | + | |
| 908 | + | |
| 909 | + | |
| 910 | + | |
| 911 | + | |
| 912 | + | |
| 913 | + | |
| 914 | + | |
| 915 | + | |
| 916 | + | |
| 917 | + | |
| 918 | + | |
| 919 | + | |
| 920 | + | |
| 921 | + | |
| 922 | + | |
| 923 | + | |
| 924 | + | |
| 925 | + | |
| 926 | + | |
| 927 | + | |
| 928 | + | |
| 929 | + | |
| 930 | + | |
| 931 | + | |
| 932 | + | |
| 933 | + | |
| 934 | + | |
| 935 | + | |
| 936 | + | |
| 937 | + | |
| 938 | + | |
| 939 | + | |
| 940 | + | |
| 941 | + | |
| 942 | + | |
| 943 | + | |
| 944 | + | |
| 945 | + | |
| 946 | + | |
| 947 | + | |
| 948 | + | |
| 949 | + | |
895 | 950 | | |
896 | 951 | | |
897 | 952 | | |
| |||
923 | 978 | | |
924 | 979 | | |
925 | 980 | | |
926 | | - | |
927 | | - | |
928 | | - | |
929 | | - | |
| 981 | + | |
| 982 | + | |
| 983 | + | |
930 | 984 | | |
931 | 985 | | |
932 | 986 | | |
933 | 987 | | |
934 | | - | |
935 | | - | |
936 | | - | |
937 | | - | |
938 | | - | |
939 | | - | |
940 | | - | |
941 | | - | |
942 | | - | |
943 | | - | |
944 | | - | |
945 | | - | |
946 | | - | |
947 | | - | |
948 | | - | |
949 | | - | |
950 | | - | |
951 | | - | |
952 | | - | |
953 | | - | |
954 | | - | |
| 988 | + | |
955 | 989 | | |
956 | 990 | | |
957 | 991 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
| 1 | + | |
| 2 | + | |
| 3 | + | |
| 4 | + | |
| 5 | + | |
| 6 | + | |
| 7 | + | |
| 8 | + | |
| 9 | + | |
| 10 | + | |
| 11 | + | |
| 12 | + | |
| 13 | + | |
| 14 | + | |
| 15 | + | |
| 16 | + | |
| 17 | + | |
| 18 | + | |
| 19 | + | |
| 20 | + | |
| 21 | + | |
| 22 | + | |
| 23 | + | |
| 24 | + | |
| 25 | + | |
| 26 | + | |
| 27 | + | |
| 28 | + | |
| 29 | + | |
| 30 | + | |
| 31 | + | |
| 32 | + | |
| 33 | + | |
| 34 | + | |
| 35 | + | |
| 36 | + | |
| 37 | + | |
| 38 | + | |
| 39 | + | |
| 40 | + | |
| 41 | + | |
| 42 | + | |
| 43 | + | |
| 44 | + | |
| 45 | + | |
| 46 | + | |
| 47 | + | |
| 48 | + | |
| 49 | + | |
| 50 | + | |
| 51 | + | |
| 52 | + | |
| 53 | + | |
| 54 | + | |
| 55 | + | |
| 56 | + | |
| 57 | + | |
| 58 | + | |
| 59 | + | |
| 60 | + | |
| 61 | + | |
| 62 | + | |
| 63 | + | |
| 64 | + | |
| 65 | + | |
| 66 | + | |
| 67 | + | |
| 68 | + | |
| 69 | + | |
| 70 | + | |
| 71 | + | |
| 72 | + | |
| 73 | + | |
| 74 | + | |
| 75 | + | |
| 76 | + | |
| 77 | + | |
| 78 | + | |
| 79 | + | |
| 80 | + | |
| 81 | + | |
| 82 | + | |
| 83 | + | |
| 84 | + | |
| 85 | + | |
| 86 | + | |
| 87 | + | |
| 88 | + | |
| 89 | + | |
| 90 | + | |
| 91 | + | |
| 92 | + | |
| 93 | + | |
| 94 | + | |
| 95 | + | |
| 96 | + | |
| 97 | + | |
| 98 | + | |
| 99 | + | |
| 100 | + | |
| 101 | + | |
| 102 | + | |
| 103 | + | |
| 104 | + | |
| 105 | + | |
| 106 | + | |
| 107 | + | |
| 108 | + | |
| 109 | + | |
| 110 | + | |
| 111 | + | |
| 112 | + | |
| 113 | + | |
| 114 | + | |
| 115 | + | |
| 116 | + | |
| 117 | + | |
| 118 | + | |
| 119 | + | |
| 120 | + | |
| 121 | + | |
| 122 | + | |
| 123 | + | |
| 124 | + | |
| 125 | + | |
| 126 | + | |
| 127 | + | |
| 128 | + | |
| 129 | + | |
| 130 | + | |
| 131 | + | |
| 132 | + | |
| 133 | + | |
| 134 | + | |
| 135 | + | |
| 136 | + | |
| 137 | + | |
| 138 | + | |
| 139 | + | |
| 140 | + | |
| 141 | + | |
| 142 | + | |
| 143 | + | |
| 144 | + | |
| 145 | + | |
| 146 | + | |
| 147 | + | |
| 148 | + | |
| 149 | + | |
| 150 | + | |
| 151 | + | |
| 152 | + | |
| 153 | + | |
| 154 | + | |
| 155 | + | |
| 156 | + | |
| 157 | + | |
| 158 | + | |
| 159 | + | |
| 160 | + | |
| 161 | + | |
| 162 | + | |
| 163 | + | |
| 164 | + | |
| 165 | + | |
| 166 | + | |
| 167 | + | |
| 168 | + | |
| 169 | + | |
| 170 | + | |
| 171 | + | |
| 172 | + | |
| 173 | + | |
| 174 | + | |
| 175 | + | |
| 176 | + | |
| 177 | + | |
| 178 | + | |
| 179 | + | |
| 180 | + | |
| 181 | + | |
| 182 | + | |
| 183 | + | |
| 184 | + | |
| 185 | + | |
| 186 | + | |
| 187 | + | |
| 188 | + | |
| 189 | + | |
| 190 | + | |
| 191 | + | |
| 192 | + | |
| 193 | + | |
| 194 | + | |
| 195 | + | |
| 196 | + | |
| 197 | + | |
| 198 | + | |
| 199 | + | |
| 200 | + | |
| 201 | + | |
| 202 | + | |
| 203 | + | |
| 204 | + | |
| 205 | + | |
| 206 | + | |
| 207 | + | |
| 208 | + | |
| 209 | + | |
| 210 | + | |
| 211 | + | |
| 212 | + | |
| 213 | + | |
| 214 | + | |
| 215 | + | |
| 216 | + | |
| 217 | + | |
| 218 | + | |
| 219 | + | |
| 220 | + | |
| 221 | + | |
| 222 | + | |
| 223 | + | |
| 224 | + | |
| 225 | + | |
| 226 | + | |
| 227 | + | |
| 228 | + | |
| 229 | + | |
| 230 | + | |
| 231 | + | |
| 232 | + | |
| 233 | + | |
| 234 | + | |
| 235 | + | |
| 236 | + | |
| 237 | + | |
| 238 | + | |
| 239 | + | |
| 240 | + | |
| 241 | + | |
| 242 | + | |
| 243 | + | |
0 commit comments