Can a mid-sized energy model get Mo–Ta right and still predict the search's alloys?
Yes. Six numbers per element did both, with 6.1 meV/atom error on the search's alloys; it became the model that ships.
In the log: the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)
confirmedDate 2026-09-23 12:3x, as written in the logrung 4 · DFT3 predictions · 1 result paragraphEXPERIMENTS.md lines 15969–15977, lines 16065–16088, lines 16190–16196
What E237 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E237.svg).
Pre-registration
(1)
e5's ensemble meets
E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9.
confirmede5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held …
(2)
e4 does not (mean > +20).
confirmedin direction, not magnitude — e4 misses the Mo–Ta target …
(3)
Throughput falls roughly as d³: e5 at 25–60 energies/s
(5–12× slower than v5), so a rung-1 walk costs 2–5 h.
mootthe exact evaluator made rung-1 cost nearly independent …
The pre-registration, as written
E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)
v5's exact recipe with --embedding 4, 5, 6, seeds 0–2 (GPU), scored like E229, plus each
model's FastECE throughput on the same 432-site cell. Predictions. (1) e5's ensemble meets
E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9.
(2) e4 does not (mean > +20). (3) Throughput falls roughly as d³: e5 at 25–60 energies/s
(5–12× slower than v5), so a rung-1 walk costs 2–5 h. Decision: rung 0/1 for the re-walk is
the smallest embedding meeting (1)'s targets; if none does below 9, rung 1 stays on v5 for the
ordering temperature and the e9 ensemble scores rung 0 only, stated as such.
E237's embedding-6 seed ensemble (old labels; E238's fallback), rung 1 as the mean of the three
members through the exact evaluator. Four shards (core budget: two GPU fits and E224's search share
the Mac). Predictions. (1) Of the 24 finds E223 left at "no transition above 100 K", at least 10
now carry a verdict. (2) The Mo–Ta-rich finds (Mo₄₂Ta₂₉Ti₂₉, Mo₄₈W₂₉Ta₁₂Nb₈, Mo₅₄Ta₁₆W₁₃Ti₇Hf₅)
order > 150 K higher than under v5 (e6 binds B2 Mo–Ta ~35 meV/atom deeper). (3) The top-p pick
changes. Rung 0's absolute formation energies carry E239's open reference question; rung 1 does
not (composition-linear). runs/e240_chain.sh.
Results
EXPERIMENTS.md · line 16065
E237 RESULT (15:0x) — embedding 5 passes its bars; embedding 6 is the one that holds both faces.
Seed ensembles (3 each), v5's recipe, scored like E229 (runs/e237_report.txt), then on E223's three
DFT random cells (scripts/dft/e223_dft_score.py, same cells as the E223 rung-4 result):
(meV/atom.) (1) CONFIRMED — e5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9).
(2) CONFIRMED in direction, not magnitude — e4 misses the Mo–Ta target at +17.3, not the
predicted > +20. (3) moot — the exact evaluator made rung-1 cost nearly independent of
embedding (the chain's own throughput step then failed on a wrong call signature; the numbers are
the exact-evaluator ones above). Decision. E237's rule alone picks e5. But the requirement
recorded at 14:1x, before these numbers existed — the model that ships must also hold the search's
own random cells — removes it: e5 is off by −20.3 on the seven-element pick and +23.9 on
Mo₇₀Hf₂₀Ti₁₀ with a seed spread of 2.9, confidently wrong the way e9 was. Embedding 6 is the
only one that holds both faces: v5's accuracy on the random cells (6.1 vs 6.3; −13.5, −4.4, −0.4)
and e9's on ordered states. Rung 0/1 go to embedding 6. Caveat stated: three random cells is a
thin test, and e5's miss could be one composition's; the final fit carries an e5 arm on the new
labels so the choice is re-tested, not assumed.
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 15969–15977
E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)
v5's exact recipe with --embedding 4, 5, 6, seeds 0–2 (GPU), scored like E229, plus each
model's FastECE throughput on the same 432-site cell. Predictions. (1) e5's ensemble meets
E229's DFT targets (four Mo–Ta cells' mean < +15, eleven-cell MAE ≤ 14) with held-48 ≤ 9.
(2) e4 does not (mean > +20). (3) Throughput falls roughly as d³: e5 at 25–60 energies/s
(5–12× slower than v5), so a rung-1 walk costs 2–5 h. Decision: rung 0/1 for the re-walk is
the smallest embedding meeting (1)'s targets; if none does below 9, rung 1 stays on v5 for the
ordering temperature and the e9 ensemble scores rung 0 only, stated as such.
EXPERIMENTS.md · lines 16065–16088
E237 RESULT (15:0x) — embedding 5 passes its bars; embedding 6 is the one that holds both faces.
Seed ensembles (3 each), v5's recipe, scored like E229 (runs/e237_report.txt), then on E223's three
DFT random cells (scripts/dft/e223_dft_score.py, same cells as the E223 rung-4 result):
(meV/atom.) (1) CONFIRMED — e5 meets every target (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9).
(2) CONFIRMED in direction, not magnitude — e4 misses the Mo–Ta target at +17.3, not the
predicted > +20. (3) moot — the exact evaluator made rung-1 cost nearly independent of
embedding (the chain's own throughput step then failed on a wrong call signature; the numbers are
the exact-evaluator ones above). Decision. E237's rule alone picks e5. But the requirement
recorded at 14:1x, before these numbers existed — the model that ships must also hold the search's
own random cells — removes it: e5 is off by −20.3 on the seven-element pick and +23.9 on
Mo₇₀Hf₂₀Ti₁₀ with a seed spread of 2.9, confidently wrong the way e9 was. Embedding 6 is the
only one that holds both faces: v5's accuracy on the random cells (6.1 vs 6.3; −13.5, −4.4, −0.4)
and e9's on ordered states. Rung 0/1 go to embedding 6. Caveat stated: three random cells is a
thin test, and e5's miss could be one composition's; the final fit carries an e5 arm on the new
labels so the choice is re-tested, not assumed.
EXPERIMENTS.md · lines 16190–16196
E237's embedding-6 seed ensemble (old labels; E238's fallback), rung 1 as the mean of the three
members through the exact evaluator. Four shards (core budget: two GPU fits and E224's search share
the Mac). Predictions. (1) Of the 24 finds E223 left at "no transition above 100 K", at least 10
now carry a verdict. (2) The Mo–Ta-rich finds (Mo₄₂Ta₂₉Ti₂₉, Mo₄₈W₂₉Ta₁₂Nb₈, Mo₅₄Ta₁₆W₁₃Ti₇Hf₅)
order > 150 K higher than under v5 (e6 binds B2 Mo–Ta ~35 meV/atom deeper). (3) The top-p pick
changes. Rung 0's absolute formation energies carry E239's open reference question; rung 1 does
not (composition-linear). runs/e240_chain.sh.
Related entries
E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered…
E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain…
E238 — the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)
E224 — has all eight arms; E225 and E226 pre-registered and queued (2026-09-22 06:2x)
E239 — pin the label frame's pure-element references with QE anchor cells (pre-registered…