Would a newer version of the screening model fix its distorted energies?
Yes. A newer, freely licensed model cut the error from 70.4 to 23.0 meV/atom and ranked the ten alloys best.
In the log: A drop-in model replacement fixes the screening potential
recordedDate 2026-09-12, as written in the logunclassified0 predictions · 1 result paragraphEXPERIMENTS.md lines 448–502
What E13 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E13.svg).
Results
EXPERIMENTS.md · line 460
Result.
The full record
EXPERIMENTS.md · lines 448–502
E13 — A drop-in model replacement fixes the screening potential
Date 2026-09-12 · Question Is the compressed dynamic range a property of
MACE-MP-0, or of foundation potentials generally? · Provenance…/rescore.py
Method. Re-score the ten alloys Quantum ESPRESSO already evaluated, at exactly
the stored unrelaxed geometries, with a second foundation model. Mixing enthalpy is
a difference within one model — alloy minus its own elemental references on the same
bcc 2x2x2 cells — so every per-element energy zero cancels and cross-model offsets
are irrelevant. No new DFT. QE dH_mix spans -58.3 to +91.0 meV/atom.
Result.
model
licence
Spearman
MAE
bias
slope
sign
MACE-MP-0 medium (current)
MIT
+0.770
70.4
-60.4
0.560
80%
MACE-MPA-0 medium
MIT
+0.927
23.0
-23.0
1.118
90%
MACE-OMAT-0 medium
ASL, academic only
+0.879
12.4
-10.6
1.032
90%
MAE and bias in meV/atom; slope is of QE on model, 1.0 being perfect.
The licence and the metric point the same way. MACE-OMAT-0 is the most accurate in
absolute terms — MAE 12.4, about six times better than the current model — but it is
academic-use-only. For a ranking task the relevant metric is rank correlation, and
there the MIT-licensed MACE-MPA-0 wins outright (0.927 against 0.879). The
commercially clean choice is also the better screening choice.
ORB v3 (Apache-2.0) was not tested: its ASE calculator has moved across three module
paths between releases and requires an atoms_adapter argument this version does not
document. Worth revisiting, since the literature reports it as the one model that
reproduces a bcc refractory miscibility gap where MACE-MP-0 fails by 109%.
Pipeline validation. The MACE-MP-0 column reproduces the book's published
unrelaxed figures — Spearman 0.77, MAE 70, bias -60 — to the digits quoted. The
re-scoring path is therefore measuring the same quantity the original runs did.
Interpretation. The compressed dynamic range is a property of the model, not of
foundation potentials in general — every newer checkpoint tested fixes it. MACE-MPA-0
(MPtrj + sAlex, 9.06M parameters, MIT licence) cuts the error threefold and brings
the slope to 1.118; MACE-OMAT-0 reaches 12.4 meV/atom and slope 1.032 but cannot be
used commercially.
The defect that made rewarding a controller with MACE dH actively harmful — a
landscape three to four times too steep — is removed by changing one model alias.
No fine-tuning, no new DFT, no cloud spend.
Caveats.
n = 10, and unrelaxed only. The published slope of 0.28 was fitted to the
relaxed-at-pressure set, whose geometries were not stored, so the two slope
figures are from different subsets and 0.560 is not a contradiction of 0.28.
The relaxed set is where sign disagreement was 5 of 10 and where the defect bites
hardest; it cannot be re-scored without regenerating those geometries.
EquiformerV3+DeNS-OAM (MIT) and SevenNet-Omni-i12 (MIT) remain untested, as does
ORB v3 for the API reason above.