Experiments · E229

Would giving the energy model more numbers per element fix its Mo–Ta error?

Partly. Error on 11 DFT test cells fell from 20.6 to 8.8 meV/atom, but unseen alloys got worse (7.1 to 11.0).

In the log: is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)

mixedDate 2026-09-22 19:4x, as written in the logrung 4 · DFT3 predictions · 2 result paragraphsEXPERIMENTS.md lines 15773–15789, lines 15844–15850, lines 15918–15925
exp E229 diagram
What E229 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E229.svg).

Pre-registration

  1. (1)
    Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to < +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2.
    confirmedon both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the …
  2. (2)
    It does not: Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random Mo–Ta frame at the label's definition; triplet cutoff).
    falsified
  3. (3)
    Whatever the means do, the e9 seed spread is reported: a model with more freedom and the same data may be less determined (spread up) — which would make the ensemble mandatory rather than advisable. Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.
    no verdict written against it
The pre-registration, as written

E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)

Every eCE since v2 compresses the nine species into a 3-dimensional embedding (16-16 network, pair 6.0 Å, triplet 4.5 Å). Thirty-six unlike pairs through a 3-d bottleneck is a plausible reason a model cannot make Mo–Ta as deep as its own B2 label while holding the quaternaries. Design: v5's exact recipe (--labels runs/rhea_labels_formation_strain0.10_pureref9.jsonl --holdout "Mo,Nb,Ta,W;Mo,Nb,Ta,V,W"), RHEA rows only, --embedding 9, seeds 0, 1, 2 — against the three v5 seeds as the embedding-3 arm. One variable. Scored with score_unbiased_qe.py as ensembles. Predictions. (1) Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to < +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2. (2) It does not: Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random Mo–Ta frame at the label's definition; triplet cutoff). (3) Whatever the means do, the e9 seed spread is reported: a model with more freedom and the same data may be less determined (spread up) — which would make the ensemble mandatory rather than advisable. Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.

Results

EXPERIMENTS.md · line 15844

E229, re-read against the literature before its result (20:2x). Müller & Natarajan (npj Comput. Mater. 2025) find 3–4 embedding dimensions optimal on a group-5/6 senary and worse extrapolation at 6: larger embeddings stop exploiting chemical trends. Our set adds group 4, the case they reserve larger embeddings for, so E229 still measures something — but prediction (2) is now the expected branch, and an e9 ensemble that is worse on the unbiased set is their finding reproduced, not a null. Their pair cutoff is 10 Å; ours, 6.0 Å, was set from a linear CE's parameter budget (l. 629) and never revisited. docs/research/2026-09-22-rung0-whole-picture.md.

EXPERIMENTS.md · line 15918

E229 RESULT. Embedding 9 (3 seeds, GPU) against embedding 3 (3 CPU seeds; 3 GPU seeds as the device control). Eleven unbiased QE cells, ensemble MAE: e3-CPU 20.6, e3-GPU 18.0, e9 8.8; B2 Mo–Ta +29.5 / +18.1 / +0.2; the four Mo–Ta cells' mean +34 / +23 / +9 (corrected 2026-09-24 00:09: the GPU control's mean is +25, from runs/e229_report.txt per-cell +18 +27 +34 +21); training RMSE ~13.6 / 12.9 / 6.8. Held-out Mo–Nb–Ta–W MAE 7.1 / 6.5 / 11.0 (seeds 20.9, 8.3, 12.3). (1) confirmed on both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the held-48 guard (≤ 8.2); (2) falsified. Capacity was binding for ordered states; nine dimensions lose the chemical trends that carry an unseen quaternary — Müller & Natarajan's Fig. 7 reproduced. The device changes nothing beyond seed noise (18.0 vs 20.6, spreads overlap).

The full record

This entry is written in 3 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 15773–15789

E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered 2026-09-22 19:4x)

Every eCE since v2 compresses the nine species into a 3-dimensional embedding (16-16 network, pair 6.0 Å, triplet 4.5 Å). Thirty-six unlike pairs through a 3-d bottleneck is a plausible reason a model cannot make Mo–Ta as deep as its own B2 label while holding the quaternaries. Design: v5's exact recipe (--labels runs/rhea_labels_formation_strain0.10_pureref9.jsonl --holdout "Mo,Nb,Ta,W;Mo,Nb,Ta,V,W"), RHEA rows only, --embedding 9, seeds 0, 1, 2 — against the three v5 seeds as the embedding-3 arm. One variable. Scored with score_unbiased_qe.py as ensembles. Predictions. (1) Capacity binds: the e9 ensemble's mean residual on the four Mo–Ta cells falls from +34 to < +15, its eleven-cell MAE from 20.6 to ≤ 14, held-48 MAE ≤ 8.2. (2) It does not: Mo–Ta mean within ±5 of +34 and eleven-cell MAE within ±3 of 20.6 → capacity is ruled out and the Mo–Ta error lives in the labels or the cluster basis (next: QE on RHEA's own random Mo–Ta frame at the label's definition; triplet cutoff). (3) Whatever the means do, the e9 seed spread is reported: a model with more freedom and the same data may be less determined (spread up) — which would make the ensemble mandatory rather than advisable. Chain runs/e229_chain.sh: one fit at a time, two threads, load-gated, yields to searches.

EXPERIMENTS.md · lines 15844–15850

E229, re-read against the literature before its result (20:2x). Müller & Natarajan (npj Comput. Mater. 2025) find 3–4 embedding dimensions optimal on a group-5/6 senary and worse extrapolation at 6: larger embeddings stop exploiting chemical trends. Our set adds group 4, the case they reserve larger embeddings for, so E229 still measures something — but prediction (2) is now the expected branch, and an e9 ensemble that is worse on the unbiased set is their finding reproduced, not a null. Their pair cutoff is 10 Å; ours, 6.0 Å, was set from a linear CE's parameter budget (l. 629) and never revisited. docs/research/2026-09-22-rung0-whole-picture.md.

EXPERIMENTS.md · lines 15918–15925

E229 RESULT. Embedding 9 (3 seeds, GPU) against embedding 3 (3 CPU seeds; 3 GPU seeds as the device control). Eleven unbiased QE cells, ensemble MAE: e3-CPU 20.6, e3-GPU 18.0, e9 8.8; B2 Mo–Ta +29.5 / +18.1 / +0.2; the four Mo–Ta cells' mean +34 / +23 / +9 (corrected 2026-09-24 00:09: the GPU control's mean is +25, from runs/e229_report.txt per-cell +18 +27 +34 +21); training RMSE ~13.6 / 12.9 / 6.8. Held-out Mo–Nb–Ta–W MAE 7.1 / 6.5 / 11.0 (seeds 20.9, 8.3, 12.3). (1) confirmed on both DFT targets (Mo–Ta < +15, MAE ≤ 14), failed on the held-48 guard (≤ 8.2); (2) falsified. Capacity was binding for ordered states; nine dimensions lose the chemical trends that carry an unseen quaternary — Müller & Natarajan's Fig. 7 reproduced. The device changes nothing beyond seed noise (18.0 vs 20.6, spreads overlap).

Built with PRISMWebsite and visualizations made using Claude