Will first-principles DFT agree with the atomistic model on which alloys are more stable?
Yes. Tested later on a repaired comparison, the two agreed to 28 meV/atom on average, inside the predicted 30.
In the log: DFT rung: prediction recorded before the first result
confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04rung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 4105–4134
What E75 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E75.svg).
Results
EXPERIMENTS.md · line 4105
E75 — DFT rung: prediction recorded before the first result
Written before the running job returns, so it can falsify something.
What the rung can and cannot be. A DFT total energy is meaningless next to a MACE one:
different pseudopotentials put them on different absolute scales, and no amount of care
makes them subtractable. So the top rung cannot recompute the verdict. What it can do is
check the one quantity the verdict rests on - the energy difference between a candidate
and the reference - computed the same way in both models:
E_DFT(candidate) - E_DFT(MoNbTaW) against E_MACE(candidate) - E_MACE(MoNbTaW)
Same cell, same occupancies, same fixed lattice, neither relaxed. All three conventions
match by construction rather than by assertion, which is the only reason this comparison is
allowed at all.
Prediction: the two agree to within 30 meV/atom. Both are PBE; MACE-MPA-0 is trained on
PBE data, so differencing two similar bcc refractory alloys should cancel most of the
systematic error, and published MACE-MP benchmarks sit at 30-50 meV/atom on formation
energies with differences between like structures doing better.
Falsification: a disagreement above 50 meV/atom means the ranking is not real. The gaps
the project is built on are 57 to 130 meV/atom, so an error of that size reorders them.
Known error term, measured, not assumed. The k-point grid is 3x3x3. The convergence
study running on this machine gives, for the 16-atom cell: 2x2x2 is 55.7 meV/atom from
3x3x3, 3x3x3 is 3.8 from 4x4x4, and 4x4x4 is 0.29 from 6x6x6. So 3x3x3 carries about
4 meV/atom. It cancels in part between two structures in the same cell on the same grid,
but not exactly, since different occupancies give different Fermi surfaces. Small against a
30 meV prediction; quoted rather than ignored.
The full record
EXPERIMENTS.md · lines 4105–4134
E75 — DFT rung: prediction recorded before the first result
Written before the running job returns, so it can falsify something.
What the rung can and cannot be. A DFT total energy is meaningless next to a MACE one:
different pseudopotentials put them on different absolute scales, and no amount of care
makes them subtractable. So the top rung cannot recompute the verdict. What it can do is
check the one quantity the verdict rests on - the energy difference between a candidate
and the reference - computed the same way in both models:
E_DFT(candidate) - E_DFT(MoNbTaW) against E_MACE(candidate) - E_MACE(MoNbTaW)
Same cell, same occupancies, same fixed lattice, neither relaxed. All three conventions
match by construction rather than by assertion, which is the only reason this comparison is
allowed at all.
Prediction: the two agree to within 30 meV/atom. Both are PBE; MACE-MPA-0 is trained on
PBE data, so differencing two similar bcc refractory alloys should cancel most of the
systematic error, and published MACE-MP benchmarks sit at 30-50 meV/atom on formation
energies with differences between like structures doing better.
Falsification: a disagreement above 50 meV/atom means the ranking is not real. The gaps
the project is built on are 57 to 130 meV/atom, so an error of that size reorders them.
Known error term, measured, not assumed. The k-point grid is 3x3x3. The convergence
study running on this machine gives, for the 16-atom cell: 2x2x2 is 55.7 meV/atom from
3x3x3, 3x3x3 is 3.8 from 4x4x4, and 4x4x4 is 0.29 from 6x6x6. So 3x3x3 carries about
4 meV/atom. It cancels in part between two structures in the same cell on the same grid,
but not exactly, since different occupancies give different Fermi surfaces. Small against a
30 meV prediction; quoted rather than ignored.