Are the training energies biased because the machine-learned potential is too soft?
No. The potential is not soft (force slope 1.014), but a volume-fitting step leaves labels 17.2 meV/atom too high.
In the log: Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)
mixedDate 2026-09-22 19:1x, as written in the logrung 4 · DFT4 predictions · 1 result paragraphEXPERIMENTS.md lines 15670–15713, lines 15715–15742
What E228 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E228.svg).
Pre-registration
(1)
Force parity F_MACE vs F_DFT (pooled components, through the origin):
slope s_F < 0.90; falsified at ≥ 0.95.
falsified** force slope s_F = **1 …
(2)
Virial parity (tr/3 per frame, with an
intercept): slope s_V < 0.90; falsified at ≥ 0.95.
falsified** virial slope s_V = **1 …
(3)
The implied label bias
B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a
uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE.
falsifiedimplied softening bias −0 …
(4)
δ > 0 on most frames
with mean ≥ 5 meV/atom; falsified if the mean is < 1.
confirmedthe stored a_relaxed is not MACE's minimum — …
The pre-registration, as written
E228 — Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)
Why now. Every label v5 was fitted on is E_DFT − [MACE(as RHEA has it) − MACE(ideal sites, MACE-relaxed a)] — the rattle-plus-strain energy removed by MACE-MPA-0, worth ~100 meV/atom
per frame (97 on the first row of the file). The docstring trusts MACE here because "E78
measured it at 1–4 meV/atom", but E78 measured something else: mixing-energy gaps
between compositions at one fixed ideal geometry. MACE on a geometry change of one
composition — the thing subtracted — has never been measured. Universal potentials are
documented to be systematically soft (PES curvature under-estimated). If MACE-MPA-0 is soft
on these frames the correction is short and every RHEA label sits high by the shortfall —
which would read exactly like the +32/+40 meV/atom v5 carries above QE on the random Mo–Ta
cells and E222's random twin, and like v7b's −7.7 on held-out RHEA (a model pulled toward
clean ideal-lattice QE labels, scored against biased ones). RHEA ships DFT forces and virials
on every frame, so this is testable with no new DFT.
Protocol. 150 frames drawn (seed 0) from the 4,326 labels v5 was fitted on
(runs/rhea_labels_formation_strain0.10_pureref9.jsonl), matched to db_9elem.xyz by index.
Per frame, MACE medium-mpa-0, float64: forces and virial at RHEA's geometry; energy at ideal
sites at RHEA's a; energy at ideal sites at the stored a_relaxed. Script
scripts/validate/mace_softening.py, checkpointed per frame.
Controls (no result is read unless both pass): the recomputed −½ΣF_DFT·u/N reproduces
the stored e_rattle_harmonic to < 0.01 meV/atom, and the recomputed correction reproduces the
stored correction_per_atom to < 0.1 meV/atom.
Predictions.(1) Force parity F_MACE vs F_DFT (pooled components, through the origin):
slope s_F < 0.90; falsified at ≥ 0.95. (2) Virial parity (tr/3 per frame, with an
intercept): slope s_V < 0.90; falsified at ≥ 0.95. (3) The implied label bias
B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a
uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE.
Decision. Both slopes ≥ 0.95 → softening is not the cause; the pre-registered per-group
VASP-vs-QE offset stands as the next refit. Either slope < 0.90 with (3) in range → the
labels are biased high by the correction itself; the fix is a DFT-calibrated correction and a
relabel (QE on three RHEA frames, as-is and ideal, pins the scale), not a per-group
offset, and v7b's "failure" was scored against the biased frame. In between → the
calibration run decides.
Addendum before the full run (19:2x). Smoke test, 3 frames: both controls pass at 0.0000
meV/atom; s_F = 1.02 on 3 frames (no conclusion from three). But MACE(ideal, a_in) −
MACE(ideal, a_relaxed) averages −10 meV/atom on them, and if a_relaxed were MACE's
minimum that difference could never be negative. The labeller takes a_relaxed as the vertex
of a parabola fitted over ±10 % in a — a range over which E(a) is strongly anharmonic. Added
measurement: δ = E_MACE(ideal, a_relaxed) − min_a E_MACE(ideal, a) by a bounded 1-D
minimisation, per frame. Every meV of δ is a meV the label sits high (it describes the ideal
lattice at a volume the configuration does not adopt). Prediction (4): δ > 0 on most frames
with mean ≥ 5 meV/atom; falsified if the mean is < 1.
Results
EXPERIMENTS.md · line 15715
E228 RESULT (2026-09-22 19:2x): MACE is not soft — falsified; the labeller's volume fit misses by +17 meV/atom — confirmed.
150 frames, both controls at 0.0000 meV/atom. (1) falsified: force slope s_F = 1.014
(cosine 0.987; harmonic-work ratio 0.998); by element Cr 0.94, Hf 0.96, Mo 1.05, Nb 1.06, Ta
1.03, Ti 0.98, V 1.00, W 1.01, Zr 1.01. (2) falsified: virial slope s_V = 1.072 (corr
0.978, p_DFT from −2.3 to +2.3 eV/atom); the +146 meV/atom intercept is an equilibrium-volume
offset worth ≈ 0.2 % in a, ≈ 0.5 meV/atom. (3) falsified: implied softening bias −0.2
meV/atom (p10–p90 −1.1…+1.1). (4) confirmed: the stored a_relaxed is not MACE's minimum —
E(ideal, a_in) < E(ideal, a_relaxed) on 128 of 150 frames; a_relaxed − a_min = +0.047 Å on
average (max 0.078); δ = +17.2 meV/atom mean (median 16.8, p90 27.6, max 35.3). The
parabola over ±10 % in a puts its vertex on the soft (expanded) side of a curve that is steep
under compression, so every alloy label sits high by δ. What this does and does not
explain. It does not explain the +32 on the 16-atom Mo–Ta cells: the table of v5's signed
residuals on the 21 QE store rows (all v7b could score) shows no offset on 54-atom random-
like cells — −4 ± 11 over 11 cells — while the four 16-atom Mo–Ta cells (E191; k 7³, v3
references, matching Widom's PBE B2 −186 and random −90…−127) sit at +23…+40. Those four rows
were in v7b's training set; pulling the Mo–Ta region down by that much is v7b's −7.7 on the
held-out Mo–Nb–Ta–W rows, without any VASP-vs-QE frame offset. The pre-registered per-group
offset is therefore the wrong fix, and the "+32/+40 on random Mo–Ta cells" this record
cited as a frame offset is a 16-atom effect (or a v2/v3 standard effect — E180, the 54-atom
comparator, is on the older 50/400 Ry, k 3³ standard; E192_MoTa_T800, 54-atom on the v3
standard, is the control and is being scored next). Also found: the store holds rows on two
DFT standards (v2: 50/400 Ry, refs k 10³; v3: 60/720 Ry, refs k 13³), each internally
consistent (alloy and references on the same cutoff and k-density), whose references differ by
+8.6 (Ta) to +42.2 (Nb) meV/atom in absolute energy — harmless per row, a precision
difference between rows. Still open from this: δ on the pure-reference frames (the part of
δ that survives into formation energies is δ_alloy − Σ xᵢ δ_pure,i), and the references'
selection — each element's is the lowest corrected label among its pure frames, 3.5 (W) to
38.4 (Ti) meV/atom below its least-corrected frame, a winner's-curse estimate.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 15670–15713
E228 — Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)
Why now. Every label v5 was fitted on is E_DFT − [MACE(as RHEA has it) − MACE(ideal sites, MACE-relaxed a)] — the rattle-plus-strain energy removed by MACE-MPA-0, worth ~100 meV/atom
per frame (97 on the first row of the file). The docstring trusts MACE here because "E78
measured it at 1–4 meV/atom", but E78 measured something else: mixing-energy gaps
between compositions at one fixed ideal geometry. MACE on a geometry change of one
composition — the thing subtracted — has never been measured. Universal potentials are
documented to be systematically soft (PES curvature under-estimated). If MACE-MPA-0 is soft
on these frames the correction is short and every RHEA label sits high by the shortfall —
which would read exactly like the +32/+40 meV/atom v5 carries above QE on the random Mo–Ta
cells and E222's random twin, and like v7b's −7.7 on held-out RHEA (a model pulled toward
clean ideal-lattice QE labels, scored against biased ones). RHEA ships DFT forces and virials
on every frame, so this is testable with no new DFT.
Protocol. 150 frames drawn (seed 0) from the 4,326 labels v5 was fitted on
(runs/rhea_labels_formation_strain0.10_pureref9.jsonl), matched to db_9elem.xyz by index.
Per frame, MACE medium-mpa-0, float64: forces and virial at RHEA's geometry; energy at ideal
sites at RHEA's a; energy at ideal sites at the stored a_relaxed. Script
scripts/validate/mace_softening.py, checkpointed per frame.
Controls (no result is read unless both pass): the recomputed −½ΣF_DFT·u/N reproduces
the stored e_rattle_harmonic to < 0.01 meV/atom, and the recomputed correction reproduces the
stored correction_per_atom to < 0.1 meV/atom.
Predictions.(1) Force parity F_MACE vs F_DFT (pooled components, through the origin):
slope s_F < 0.90; falsified at ≥ 0.95. (2) Virial parity (tr/3 per frame, with an
intercept): slope s_V < 0.90; falsified at ≥ 0.95. (3) The implied label bias
B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a
uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE.
Decision. Both slopes ≥ 0.95 → softening is not the cause; the pre-registered per-group
VASP-vs-QE offset stands as the next refit. Either slope < 0.90 with (3) in range → the
labels are biased high by the correction itself; the fix is a DFT-calibrated correction and a
relabel (QE on three RHEA frames, as-is and ideal, pins the scale), not a per-group
offset, and v7b's "failure" was scored against the biased frame. In between → the
calibration run decides.
Addendum before the full run (19:2x). Smoke test, 3 frames: both controls pass at 0.0000
meV/atom; s_F = 1.02 on 3 frames (no conclusion from three). But MACE(ideal, a_in) −
MACE(ideal, a_relaxed) averages −10 meV/atom on them, and if a_relaxed were MACE's
minimum that difference could never be negative. The labeller takes a_relaxed as the vertex
of a parabola fitted over ±10 % in a — a range over which E(a) is strongly anharmonic. Added
measurement: δ = E_MACE(ideal, a_relaxed) − min_a E_MACE(ideal, a) by a bounded 1-D
minimisation, per frame. Every meV of δ is a meV the label sits high (it describes the ideal
lattice at a volume the configuration does not adopt). Prediction (4): δ > 0 on most frames
with mean ≥ 5 meV/atom; falsified if the mean is < 1.
EXPERIMENTS.md · lines 15715–15742
E228 RESULT (2026-09-22 19:2x): MACE is not soft — falsified; the labeller's volume fit misses by +17 meV/atom — confirmed.
150 frames, both controls at 0.0000 meV/atom. (1) falsified: force slope s_F = 1.014
(cosine 0.987; harmonic-work ratio 0.998); by element Cr 0.94, Hf 0.96, Mo 1.05, Nb 1.06, Ta
1.03, Ti 0.98, V 1.00, W 1.01, Zr 1.01. (2) falsified: virial slope s_V = 1.072 (corr
0.978, p_DFT from −2.3 to +2.3 eV/atom); the +146 meV/atom intercept is an equilibrium-volume
offset worth ≈ 0.2 % in a, ≈ 0.5 meV/atom. (3) falsified: implied softening bias −0.2
meV/atom (p10–p90 −1.1…+1.1). (4) confirmed: the stored a_relaxed is not MACE's minimum —
E(ideal, a_in) < E(ideal, a_relaxed) on 128 of 150 frames; a_relaxed − a_min = +0.047 Å on
average (max 0.078); δ = +17.2 meV/atom mean (median 16.8, p90 27.6, max 35.3). The
parabola over ±10 % in a puts its vertex on the soft (expanded) side of a curve that is steep
under compression, so every alloy label sits high by δ. What this does and does not
explain. It does not explain the +32 on the 16-atom Mo–Ta cells: the table of v5's signed
residuals on the 21 QE store rows (all v7b could score) shows no offset on 54-atom random-
like cells — −4 ± 11 over 11 cells — while the four 16-atom Mo–Ta cells (E191; k 7³, v3
references, matching Widom's PBE B2 −186 and random −90…−127) sit at +23…+40. Those four rows
were in v7b's training set; pulling the Mo–Ta region down by that much is v7b's −7.7 on the
held-out Mo–Nb–Ta–W rows, without any VASP-vs-QE frame offset. The pre-registered per-group
offset is therefore the wrong fix, and the "+32/+40 on random Mo–Ta cells" this record
cited as a frame offset is a 16-atom effect (or a v2/v3 standard effect — E180, the 54-atom
comparator, is on the older 50/400 Ry, k 3³ standard; E192_MoTa_T800, 54-atom on the v3
standard, is the control and is being scored next). Also found: the store holds rows on two
DFT standards (v2: 50/400 Ry, refs k 10³; v3: 60/720 Ry, refs k 13³), each internally
consistent (alloy and references on the same cutoff and k-density), whose references differ by
+8.6 (Ta) to +42.2 (Nb) meV/atom in absolute energy — harmless per row, a precision
difference between rows. Still open from this: δ on the pure-reference frames (the part of
δ that survives into formation energies is δ_alloy − Σ xᵢ δ_pure,i), and the references'
selection — each element's is the lowest corrected label among its pure frames, 3.5 (W) to
38.4 (Ti) meV/atom below its least-corrected frame, a winner's-curse estimate.
Related entries
E78 — First principles agrees with the potential about the ranking, to 1 and 4 meV/atom