Would giving the energy model more numbers per element fix its Mo–Ta error?
Partly. Error on 11 DFT test cells fell from 20.6 to 8.8 meV/atom, but unseen alloys got worse (7.1 to 11.0).
In the log: is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)
mixedDate 2026-09-22 19:4x, as written in the logrung 4 · DFT3 predictions · 2 result paragraphsEXPERIMENTS.md lines 15773–15789, lines 15844–15850, lines 15918–15925
What E229 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E229.svg).
Pre-registration
(1)
Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to
< +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2.
confirmedon both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the …
(2)
It does not:
Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out
and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random
Mo–Ta frame at the label's definition; triplet cutoff).
falsified
(3)
Whatever the means do, the
e9 seed spread is reported: a model with more freedom and the same data may be less
determined (spread up) — which would make the ensemble mandatory rather than advisable.
Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.
no verdict written against it
The pre-registration, as written
E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)
Every eCE since v2 compresses the nine species into a 3-dimensional embedding (16-16
network, pair 6.0 Å, triplet 4.5 Å). Thirty-six unlike pairs through a 3-d bottleneck is a
plausible reason a model cannot make Mo–Ta as deep as its own B2 label while holding the
quaternaries. Design: v5's exact recipe (--labels runs/rhea_labels_formation_strain0.10_pureref9.jsonl --holdout "Mo,Nb,Ta,W;Mo,Nb,Ta,V,W"), RHEA
rows only, --embedding 9, seeds 0, 1, 2 — against the three v5 seeds as the embedding-3
arm. One variable. Scored with score_unbiased_qe.py as ensembles. Predictions. (1)
Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to
< +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2. (2) It does not:
Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out
and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random
Mo–Ta frame at the label's definition; triplet cutoff). (3) Whatever the means do, the
e9 seed spread is reported: a model with more freedom and the same data may be less
determined (spread up) — which would make the ensemble mandatory rather than advisable.
Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.
Results
EXPERIMENTS.md · line 15844
E229, re-read against the literature before its result (20:2x). Müller & Natarajan (npj
Comput. Mater. 2025) find 3–4 embedding dimensions optimal on a group-5/6 senary and worse
extrapolation at 6: larger embeddings stop exploiting chemical trends. Our set adds group 4, the
case they reserve larger embeddings for, so E229 still measures something — but prediction (2)
is now the expected branch, and an e9 ensemble that is worse on the unbiased set is their
finding reproduced, not a null. Their pair cutoff is 10 Å; ours, 6.0 Å, was set from a
linear CE's parameter budget (l. 629) and never revisited. docs/research/2026-09-22-rung0-whole-picture.md.
EXPERIMENTS.md · line 15918
E229 RESULT. Embedding 9 (3 seeds, GPU) against embedding 3 (3 CPU seeds; 3 GPU seeds as the
device control). Eleven unbiased QE cells, ensemble MAE: e3-CPU 20.6, e3-GPU 18.0, e9
8.8; B2 Mo–Ta +29.5 / +18.1 / +0.2; the four Mo–Ta cells' mean +34 / +23 / +9(corrected 2026-09-24 00:09: the GPU control's mean is +25, from runs/e229_report.txt per-cell +18 +27 +34 +21); training
RMSE ~13.6 / 12.9 / 6.8. Held-out Mo–Nb–Ta–W MAE 7.1 / 6.5 / 11.0 (seeds 20.9, 8.3, 12.3).
(1) confirmed on both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the held-48 guard (≤ 8.2);
(2) falsified. Capacity was binding for ordered states; nine dimensions lose the chemical
trends that carry an unseen quaternary — Müller & Natarajan's Fig. 7 reproduced. The device
changes nothing beyond seed noise (18.0 vs 20.6, spreads overlap).
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 15773–15789
E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)
Every eCE since v2 compresses the nine species into a 3-dimensional embedding (16-16
network, pair 6.0 Å, triplet 4.5 Å). Thirty-six unlike pairs through a 3-d bottleneck is a
plausible reason a model cannot make Mo–Ta as deep as its own B2 label while holding the
quaternaries. Design: v5's exact recipe (--labels runs/rhea_labels_formation_strain0.10_pureref9.jsonl --holdout "Mo,Nb,Ta,W;Mo,Nb,Ta,V,W"), RHEA
rows only, --embedding 9, seeds 0, 1, 2 — against the three v5 seeds as the embedding-3
arm. One variable. Scored with score_unbiased_qe.py as ensembles. Predictions. (1)
Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to
< +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2. (2) It does not:
Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out
and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random
Mo–Ta frame at the label's definition; triplet cutoff). (3) Whatever the means do, the
e9 seed spread is reported: a model with more freedom and the same data may be less
determined (spread up) — which would make the ensemble mandatory rather than advisable.
Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.
EXPERIMENTS.md · lines 15844–15850
E229, re-read against the literature before its result (20:2x). Müller & Natarajan (npj
Comput. Mater. 2025) find 3–4 embedding dimensions optimal on a group-5/6 senary and worse
extrapolation at 6: larger embeddings stop exploiting chemical trends. Our set adds group 4, the
case they reserve larger embeddings for, so E229 still measures something — but prediction (2)
is now the expected branch, and an e9 ensemble that is worse on the unbiased set is their
finding reproduced, not a null. Their pair cutoff is 10 Å; ours, 6.0 Å, was set from a
linear CE's parameter budget (l. 629) and never revisited. docs/research/2026-09-22-rung0-whole-picture.md.
EXPERIMENTS.md · lines 15918–15925
E229 RESULT. Embedding 9 (3 seeds, GPU) against embedding 3 (3 CPU seeds; 3 GPU seeds as the
device control). Eleven unbiased QE cells, ensemble MAE: e3-CPU 20.6, e3-GPU 18.0, e9
8.8; B2 Mo–Ta +29.5 / +18.1 / +0.2; the four Mo–Ta cells' mean +34 / +23 / +9(corrected 2026-09-24 00:09: the GPU control's mean is +25, from runs/e229_report.txt per-cell +18 +27 +34 +21); training
RMSE ~13.6 / 12.9 / 6.8. Held-out Mo–Nb–Ta–W MAE 7.1 / 6.5 / 11.0 (seeds 20.9, 8.3, 12.3).
(1) confirmed on both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the held-48 guard (≤ 8.2);
(2) falsified. Capacity was binding for ordered states; nine dimensions lose the chemical
trends that carry an unseen quaternary — Müller & Natarajan's Fig. 7 reproduced. The device
changes nothing beyond seed noise (18.0 vs 20.6, spreads overlap).