Experiments · E237

Can a mid-sized energy model get Mo–Ta right and still predict the search's alloys?

Yes. Six numbers per element did both, with 6.1 meV/atom error on the search's alloys; it became the model that ships.

In the log: the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)

confirmedDate 2026-09-23 12:3x, as written in the logrung 4 · DFT3 predictions · 1 result paragraphEXPERIMENTS.md lines 15969–15977, lines 16065–16088, lines 16190–16196
exp E237 diagram
What E237 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E237.svg).

Pre-registration

  1. (1)
    e5's ensemble meets E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9.
    confirmede5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held …
  2. (2)
    e4 does not (mean > +20).
    confirmedin direction, not magnitude — e4 misses the Mo–Ta target …
  3. (3)
    Throughput falls roughly as d³: e5 at 25–60 energies/s (5–12× slower than v5), so a rung-1 walk costs 2–5 h.
    mootthe exact evaluator made rung-1 cost nearly independent …
The pre-registration, as written

E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)

v5's exact recipe with --embedding 4, 5, 6, seeds 0–2 (GPU), scored like E229, plus each model's FastECE throughput on the same 432-site cell. Predictions. (1) e5's ensemble meets E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9. (2) e4 does not (mean > +20). (3) Throughput falls roughly as d³: e5 at 25–60 energies/s (5–12× slower than v5), so a rung-1 walk costs 2–5 h. Decision: rung 0/1 for the re-walk is the smallest embedding meeting (1)'s targets; if none does below 9, rung 1 stays on v5 for the ordering temperature and the e9 ensemble scores rung 0 only, stated as such.

E237's embedding-6 seed ensemble (old labels; E238's fallback), rung 1 as the mean of the three members through the exact evaluator. Four shards (core budget: two GPU fits and E224's search share the Mac). Predictions. (1) Of the 24 finds E223 left at "no transition above 100 K", at least 10 now carry a verdict. (2) The Mo–Ta-rich finds (Mo₄₂Ta₂₉Ti₂₉, Mo₄₈W₂₉Ta₁₂Nb₈, Mo₅₄Ta₁₆W₁₃Ti₇Hf₅) order > 150 K higher than under v5 (e6 binds B2 Mo–Ta ~35 meV/atom deeper). (3) The top-p pick changes. Rung 0's absolute formation energies carry E239's open reference question; rung 1 does not (composition-linear). runs/e240_chain.sh.

Results

EXPERIMENTS.md · line 16065

E237 RESULT (15:0x) — embedding 5 passes its bars; embedding 6 is the one that holds both faces. Seed ensembles (3 each), v5's recipe, scored like E229 (runs/e237_report.txt), then on E223's three DFT random cells (scripts/dft/e223_dft_score.py, same cells as the E223 rung-4 result):

embedding eleven QE cells MAE four Mo–Ta cells, mean B2 Mo–Ta held-48 seed sd E223 random cells, mean abs. error
3 (v5) 20.6 +34 +29.5 7.1 8.5 6.3
4 13.3 +17.3 +6.3 8.4 10.1 10.0
5 7.7 +7.0 +4.1 6.7 4.2 16.6
6 7.5 +5.5 −5.2 7.9 4.0 6.1
9 (E229) 8.8 — +0.2 11.0 — 9.4

(meV/atom.) (1) CONFIRMED — e5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9). (2) CONFIRMED in direction, not magnitude — e4 misses the Mo–Ta target at +17.3, not the predicted > +20. (3) moot — the exact evaluator made rung-1 cost nearly independent of embedding (the chain's own throughput step then failed on a wrong call signature; the numbers are the exact-evaluator ones above). Decision. E237's rule alone picks e5. But the requirement recorded at 14:1x, before these numbers existed — the model that ships must also hold the search's own random cells — removes it: e5 is off by −20.3 on the seven-element pick and +23.9 on Mo₇₀Hf₂₀Ti₁₀ with a seed spread of 2.9, confidently wrong the way e9 was. Embedding 6 is the only one that holds both faces: v5's accuracy on the random cells (6.1 vs 6.3; −13.5, −4.4, −0.4) and e9's on ordered states. Rung 0/1 go to embedding 6. Caveat stated: three random cells is a thin test, and e5's miss could be one composition's; the final fit carries an e5 arm on the new labels so the choice is re-tested, not assumed.

The full record

This entry is written in 3 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 15969–15977

E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)

v5's exact recipe with --embedding 4, 5, 6, seeds 0–2 (GPU), scored like E229, plus each model's FastECE throughput on the same 432-site cell. Predictions. (1) e5's ensemble meets E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9. (2) e4 does not (mean > +20). (3) Throughput falls roughly as d³: e5 at 25–60 energies/s (5–12× slower than v5), so a rung-1 walk costs 2–5 h. Decision: rung 0/1 for the re-walk is the smallest embedding meeting (1)'s targets; if none does below 9, rung 1 stays on v5 for the ordering temperature and the e9 ensemble scores rung 0 only, stated as such.

EXPERIMENTS.md · lines 16065–16088

E237 RESULT (15:0x) — embedding 5 passes its bars; embedding 6 is the one that holds both faces. Seed ensembles (3 each), v5's recipe, scored like E229 (runs/e237_report.txt), then on E223's three DFT random cells (scripts/dft/e223_dft_score.py, same cells as the E223 rung-4 result):

embedding eleven QE cells MAE four Mo–Ta cells, mean B2 Mo–Ta held-48 seed sd E223 random cells, mean abs. error
3 (v5) 20.6 +34 +29.5 7.1 8.5 6.3
4 13.3 +17.3 +6.3 8.4 10.1 10.0
5 7.7 +7.0 +4.1 6.7 4.2 16.6
6 7.5 +5.5 −5.2 7.9 4.0 6.1
9 (E229) 8.8 — +0.2 11.0 — 9.4

(meV/atom.) (1) CONFIRMED — e5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9). (2) CONFIRMED in direction, not magnitude — e4 misses the Mo–Ta target at +17.3, not the predicted > +20. (3) moot — the exact evaluator made rung-1 cost nearly independent of embedding (the chain's own throughput step then failed on a wrong call signature; the numbers are the exact-evaluator ones above). Decision. E237's rule alone picks e5. But the requirement recorded at 14:1x, before these numbers existed — the model that ships must also hold the search's own random cells — removes it: e5 is off by −20.3 on the seven-element pick and +23.9 on Mo₇₀Hf₂₀Ti₁₀ with a seed spread of 2.9, confidently wrong the way e9 was. Embedding 6 is the only one that holds both faces: v5's accuracy on the random cells (6.1 vs 6.3; −13.5, −4.4, −0.4) and e9's on ordered states. Rung 0/1 go to embedding 6. Caveat stated: three random cells is a thin test, and e5's miss could be one composition's; the final fit carries an e5 arm on the new labels so the choice is re-tested, not assumed.

EXPERIMENTS.md · lines 16190–16196

E237's embedding-6 seed ensemble (old labels; E238's fallback), rung 1 as the mean of the three members through the exact evaluator. Four shards (core budget: two GPU fits and E224's search share the Mac). Predictions. (1) Of the 24 finds E223 left at "no transition above 100 K", at least 10 now carry a verdict. (2) The Mo–Ta-rich finds (Mo₄₂Ta₂₉Ti₂₉, Mo₄₈W₂₉Ta₁₂Nb₈, Mo₅₄Ta₁₆W₁₃Ti₇Hf₅) order > 150 K higher than under v5 (e6 binds B2 Mo–Ta ~35 meV/atom deeper). (3) The top-p pick changes. Rung 0's absolute formation energies carry E239's open reference question; rung 1 does not (composition-linear). runs/e240_chain.sh.

Related entries

Built with PRISMWebsite and visualizations made using Claude