Experiments · E13

Would a newer version of the screening model fix its distorted energies?

Yes. A newer, freely licensed model cut the error from 70.4 to 23.0 meV/atom and ranked the ten alloys best.

In the log: A drop-in model replacement fixes the screening potential

recordedDate 2026-09-12, as written in the logunclassified0 predictions · 1 result paragraphEXPERIMENTS.md lines 448–502
exp E13 diagram
What E13 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E13.svg).

Results

EXPERIMENTS.md · line 460

Result.

The full record

EXPERIMENTS.md · lines 448–502

E13 — A drop-in model replacement fixes the screening potential

Date 2026-09-12 · Question Is the compressed dynamic range a property of MACE-MP-0, or of foundation potentials generally? · Provenance …/rescore.py

Method. Re-score the ten alloys Quantum ESPRESSO already evaluated, at exactly the stored unrelaxed geometries, with a second foundation model. Mixing enthalpy is a difference within one model — alloy minus its own elemental references on the same bcc 2x2x2 cells — so every per-element energy zero cancels and cross-model offsets are irrelevant. No new DFT. QE dH_mix spans -58.3 to +91.0 meV/atom.

Result.

model licence Spearman MAE bias slope sign
MACE-MP-0 medium (current) MIT +0.770 70.4 -60.4 0.560 80%
MACE-MPA-0 medium MIT +0.927 23.0 -23.0 1.118 90%
MACE-OMAT-0 medium ASL, academic only +0.879 12.4 -10.6 1.032 90%

MAE and bias in meV/atom; slope is of QE on model, 1.0 being perfect.

The licence and the metric point the same way. MACE-OMAT-0 is the most accurate in absolute terms — MAE 12.4, about six times better than the current model — but it is academic-use-only. For a ranking task the relevant metric is rank correlation, and there the MIT-licensed MACE-MPA-0 wins outright (0.927 against 0.879). The commercially clean choice is also the better screening choice.

ORB v3 (Apache-2.0) was not tested: its ASE calculator has moved across three module paths between releases and requires an atoms_adapter argument this version does not document. Worth revisiting, since the literature reports it as the one model that reproduces a bcc refractory miscibility gap where MACE-MP-0 fails by 109%.

Pipeline validation. The MACE-MP-0 column reproduces the book's published unrelaxed figures — Spearman 0.77, MAE 70, bias -60 — to the digits quoted. The re-scoring path is therefore measuring the same quantity the original runs did.

Interpretation. The compressed dynamic range is a property of the model, not of foundation potentials in general — every newer checkpoint tested fixes it. MACE-MPA-0 (MPtrj + sAlex, 9.06M parameters, MIT licence) cuts the error threefold and brings the slope to 1.118; MACE-OMAT-0 reaches 12.4 meV/atom and slope 1.032 but cannot be used commercially. The defect that made rewarding a controller with MACE dH actively harmful — a landscape three to four times too steep — is removed by changing one model alias.

No fine-tuning, no new DFT, no cloud spend.

Caveats.

  • n = 10, and unrelaxed only. The published slope of 0.28 was fitted to the relaxed-at-pressure set, whose geometries were not stored, so the two slope figures are from different subsets and 0.560 is not a contradiction of 0.28.
  • The relaxed set is where sign disagreement was 5 of 10 and where the defect bites hardest; it cannot be re-scored without regenerating those geometries.
  • EquiformerV3+DeNS-OAM (MIT) and SevenNet-Omni-i12 (MIT) remain untested, as does ORB v3 for the API reason above.
Built with PRISMWebsite and visualizations made using Claude