Does first-principles physics rank the alloys the same way the fast potential does?
Yes. The fast potential over-binds each alloy by about 30 meV/atom, but the differences between alloys agree to within 4 meV/atom.
In the log: First principles agrees with the potential about the ranking, to 1 and 4 meV/atom
confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04unclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 4229–4274
What E78 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E78.svg).
Results
No result paragraph for this entry was found in the log.
The full record
EXPERIMENTS.md · lines 4229–4274
E78 — First principles agrees with the potential about the ranking, to 1 and 4 meV/atom
The E75 prediction, tested on the repaired comparison of E77. Seven structures at 16 atoms:
four elemental references and three alloys, every one on the same cell, the same lattice
constant, the same cutoffs, the same 3x3x3 grid and the same pseudopotentials, none relaxed.
Mixing energies, each model against its own elemental references:
composition
MACE
DFT
offset
MoNbTaW
-146
-116
-30
Mo0.62 Ta0.38
-181
-152
-29
Ta0.39 Mo0.34 W0.18 Nb0.08
-169
-143
-26
Prediction was agreement within 30 meV/atom, falsification above 50. Mean 28, worst 30.
It holds.
And the useful part is that the disagreement is an offset, not scatter. MACE over-binds
every one of these alloys by 26 to 30 meV/atom against PBE-PAW, which is the ordinary
behaviour of a foundation potential on formation energies and is nearly constant here. It
therefore cancels in the only quantity the search uses - the gap between one alloy and
another:
composition
MACE gap
DFT gap
disagreement
Mo0.62 Ta0.38 vs MoNbTaW
-35
-36
1
Ta0.39 Mo0.34 W0.18 Nb0.08 vs MoNbTaW
-23
-27
4
One and four meV/atom, against a k-point error of about four. The agreement is as close as
this measurement can resolve, both models order the three alloys identically, and the
ranking the generator learned is confirmed from first principles on the three compositions
tested.
What this does not say. Three structures, one arrangement each, unrelaxed, at 16 atoms.
It does not check the off-lattice hull, where the competitors are Laves phases nobody has
computed here in DFT, and it does not check the ordering temperature, which is a Monte Carlo
result the potential is used thousands of times inside. It checks the one number the
generator is rewarded on, and that number survives.
A second rounding trap, found inside the repair. The mixing energy must be weighted by
the cell's ACTUAL composition, not the requested one: a 16-atom cell cannot hold Mo0.6216,
it holds ten Mo and six Ta. That third of a per cent multiplies the difference between two
pseudopotential zero-points - Mo at -4951 eV/atom, Ta at -9893 - and arrives as 25 eV/atom.
With requested fractions the table read +16,677 and +37,649 meV/atom, and the equiatomic
reference, the only composition a 16-atom cell represents exactly, was the one row that
looked believable. A table in which the control is fine and everything else is absurd is
the signature of this class of bug, not of a physics result.
Related entries
E75 — DFT rung: prediction recorded before the first result
E77 — The DFT comparison was subtracting different atoms, and is withdrawn