Experiments · E228

Are the training energies biased because the machine-learned potential is too soft?

No. The potential is not soft (force slope 1.014), but a volume-fitting step leaves labels 17.2 meV/atom too high.

In the log: Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)

mixedDate 2026-09-22 19:1x, as written in the logrung 4 · DFT4 predictions · 1 result paragraphEXPERIMENTS.md lines 15670–15713, lines 15715–15742
exp E228 diagram
What E228 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E228.svg).

Pre-registration

  1. (1)
    Force parity F_MACE vs F_DFT (pooled components, through the origin): slope s_F < 0.90; falsified at ≥ 0.95.
    falsified** force slope s_F = **1 …
  2. (2)
    Virial parity (tr/3 per frame, with an intercept): slope s_V < 0.90; falsified at ≥ 0.95.
    falsified** virial slope s_V = **1 …
  3. (3)
    The implied label bias B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE.
    falsifiedimplied softening bias −0 …
  4. (4)
    δ > 0 on most frames with mean ≥ 5 meV/atom; falsified if the mean is < 1.
    confirmedthe stored a_relaxed is not MACE's minimum — …
The pre-registration, as written

E228 — Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)

Why now. Every label v5 was fitted on is E_DFT − [MACE(as RHEA has it) − MACE(ideal sites, MACE-relaxed a)] — the rattle-plus-strain energy removed by MACE-MPA-0, worth ~100 meV/atom per frame (97 on the first row of the file). The docstring trusts MACE here because "E78 measured it at 1–4 meV/atom", but E78 measured something else: mixing-energy gaps between compositions at one fixed ideal geometry. MACE on a geometry change of one composition — the thing subtracted — has never been measured. Universal potentials are documented to be systematically soft (PES curvature under-estimated). If MACE-MPA-0 is soft on these frames the correction is short and every RHEA label sits high by the shortfall — which would read exactly like the +32/+40 meV/atom v5 carries above QE on the random Mo–Ta cells and E222's random twin, and like v7b's −7.7 on held-out RHEA (a model pulled toward clean ideal-lattice QE labels, scored against biased ones). RHEA ships DFT forces and virials on every frame, so this is testable with no new DFT.

Protocol. 150 frames drawn (seed 0) from the 4,326 labels v5 was fitted on (runs/rhea_labels_formation_strain0.10_pureref9.jsonl), matched to db_9elem.xyz by index. Per frame, MACE medium-mpa-0, float64: forces and virial at RHEA's geometry; energy at ideal sites at RHEA's a; energy at ideal sites at the stored a_relaxed. Script scripts/validate/mace_softening.py, checkpointed per frame. Controls (no result is read unless both pass): the recomputed −½ΣF_DFT·u/N reproduces the stored e_rattle_harmonic to < 0.01 meV/atom, and the recomputed correction reproduces the stored correction_per_atom to < 0.1 meV/atom.

Predictions. (1) Force parity F_MACE vs F_DFT (pooled components, through the origin): slope s_F < 0.90; falsified at ≥ 0.95. (2) Virial parity (tr/3 per frame, with an intercept): slope s_V < 0.90; falsified at ≥ 0.95. (3) The implied label bias B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE. Decision. Both slopes ≥ 0.95 → softening is not the cause; the pre-registered per-group VASP-vs-QE offset stands as the next refit. Either slope < 0.90 with (3) in range → the labels are biased high by the correction itself; the fix is a DFT-calibrated correction and a relabel (QE on three RHEA frames, as-is and ideal, pins the scale), not a per-group offset, and v7b's "failure" was scored against the biased frame. In between → the calibration run decides. Addendum before the full run (19:2x). Smoke test, 3 frames: both controls pass at 0.0000 meV/atom; s_F = 1.02 on 3 frames (no conclusion from three). But MACE(ideal, a_in) − MACE(ideal, a_relaxed) averages −10 meV/atom on them, and if a_relaxed were MACE's minimum that difference could never be negative. The labeller takes a_relaxed as the vertex of a parabola fitted over ±10 % in a — a range over which E(a) is strongly anharmonic. Added measurement: δ = E_MACE(ideal, a_relaxed) − min_a E_MACE(ideal, a) by a bounded 1-D minimisation, per frame. Every meV of δ is a meV the label sits high (it describes the ideal lattice at a volume the configuration does not adopt). Prediction (4): δ > 0 on most frames with mean ≥ 5 meV/atom; falsified if the mean is < 1.

Results

EXPERIMENTS.md · line 15715

E228 RESULT (2026-09-22 19:2x): MACE is not soft — falsified; the labeller's volume fit misses by +17 meV/atom — confirmed. 150 frames, both controls at 0.0000 meV/atom. (1) falsified: force slope s_F = 1.014 (cosine 0.987; harmonic-work ratio 0.998); by element Cr 0.94, Hf 0.96, Mo 1.05, Nb 1.06, Ta 1.03, Ti 0.98, V 1.00, W 1.01, Zr 1.01. (2) falsified: virial slope s_V = 1.072 (corr 0.978, p_DFT from −2.3 to +2.3 eV/atom); the +146 meV/atom intercept is an equilibrium-volume offset worth ≈ 0.2 % in a, ≈ 0.5 meV/atom. (3) falsified: implied softening bias −0.2 meV/atom (p10–p90 −1.1…+1.1). (4) confirmed: the stored a_relaxed is not MACE's minimum — E(ideal, a_in) < E(ideal, a_relaxed) on 128 of 150 frames; a_relaxed − a_min = +0.047 Å on average (max 0.078); δ = +17.2 meV/atom mean (median 16.8, p90 27.6, max 35.3). The parabola over ±10 % in a puts its vertex on the soft (expanded) side of a curve that is steep under compression, so every alloy label sits high by δ. What this does and does not explain. It does not explain the +32 on the 16-atom Mo–Ta cells: the table of v5's signed residuals on the 21 QE store rows (all v7b could score) shows no offset on 54-atom random- like cells — −4 ± 11 over 11 cells — while the four 16-atom Mo–Ta cells (E191; k 7³, v3 references, matching Widom's PBE B2 −186 and random −90…−127) sit at +23…+40. Those four rows were in v7b's training set; pulling the Mo–Ta region down by that much is v7b's −7.7 on the held-out Mo–Nb–Ta–W rows, without any VASP-vs-QE frame offset. The pre-registered per-group offset is therefore the wrong fix, and the "+32/+40 on random Mo–Ta cells" this record cited as a frame offset is a 16-atom effect (or a v2/v3 standard effect — E180, the 54-atom comparator, is on the older 50/400 Ry, k 3³ standard; E192_MoTa_T800, 54-atom on the v3 standard, is the control and is being scored next). Also found: the store holds rows on two DFT standards (v2: 50/400 Ry, refs k 10³; v3: 60/720 Ry, refs k 13³), each internally consistent (alloy and references on the same cutoff and k-density), whose references differ by +8.6 (Ta) to +42.2 (Nb) meV/atom in absolute energy — harmless per row, a precision difference between rows. Still open from this: δ on the pure-reference frames (the part of δ that survives into formation energies is δ_alloy − Σ xᵢ δ_pure,i), and the references' selection — each element's is the lowest corrected label among its pure frames, 3.5 (W) to 38.4 (Ti) meV/atom below its least-corrected frame, a winner's-curse estimate.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 15670–15713

E228 — Is the correction under every RHEA label too soft? (pre-registered 2026-09-22 19:1x)

Why now. Every label v5 was fitted on is E_DFT − [MACE(as RHEA has it) − MACE(ideal sites, MACE-relaxed a)] — the rattle-plus-strain energy removed by MACE-MPA-0, worth ~100 meV/atom per frame (97 on the first row of the file). The docstring trusts MACE here because "E78 measured it at 1–4 meV/atom", but E78 measured something else: mixing-energy gaps between compositions at one fixed ideal geometry. MACE on a geometry change of one composition — the thing subtracted — has never been measured. Universal potentials are documented to be systematically soft (PES curvature under-estimated). If MACE-MPA-0 is soft on these frames the correction is short and every RHEA label sits high by the shortfall — which would read exactly like the +32/+40 meV/atom v5 carries above QE on the random Mo–Ta cells and E222's random twin, and like v7b's −7.7 on held-out RHEA (a model pulled toward clean ideal-lattice QE labels, scored against biased ones). RHEA ships DFT forces and virials on every frame, so this is testable with no new DFT.

Protocol. 150 frames drawn (seed 0) from the 4,326 labels v5 was fitted on (runs/rhea_labels_formation_strain0.10_pureref9.jsonl), matched to db_9elem.xyz by index. Per frame, MACE medium-mpa-0, float64: forces and virial at RHEA's geometry; energy at ideal sites at RHEA's a; energy at ideal sites at the stored a_relaxed. Script scripts/validate/mace_softening.py, checkpointed per frame. Controls (no result is read unless both pass): the recomputed −½ΣF_DFT·u/N reproduces the stored e_rattle_harmonic to < 0.01 meV/atom, and the recomputed correction reproduces the stored correction_per_atom to < 0.1 meV/atom.

Predictions. (1) Force parity F_MACE vs F_DFT (pooled components, through the origin): slope s_F < 0.90; falsified at ≥ 0.95. (2) Virial parity (tr/3 per frame, with an intercept): slope s_V < 0.90; falsified at ≥ 0.95. (3) The implied label bias B = disp_MACE·(1/s_F − 1) + strain_MACE·(1/s_V − 1), per frame (first order: softening as a uniform curvature scale), averages +15 to +45 meV/atom — the size of v5's excess over QE. Decision. Both slopes ≥ 0.95 → softening is not the cause; the pre-registered per-group VASP-vs-QE offset stands as the next refit. Either slope < 0.90 with (3) in range → the labels are biased high by the correction itself; the fix is a DFT-calibrated correction and a relabel (QE on three RHEA frames, as-is and ideal, pins the scale), not a per-group offset, and v7b's "failure" was scored against the biased frame. In between → the calibration run decides. Addendum before the full run (19:2x). Smoke test, 3 frames: both controls pass at 0.0000 meV/atom; s_F = 1.02 on 3 frames (no conclusion from three). But MACE(ideal, a_in) − MACE(ideal, a_relaxed) averages −10 meV/atom on them, and if a_relaxed were MACE's minimum that difference could never be negative. The labeller takes a_relaxed as the vertex of a parabola fitted over ±10 % in a — a range over which E(a) is strongly anharmonic. Added measurement: δ = E_MACE(ideal, a_relaxed) − min_a E_MACE(ideal, a) by a bounded 1-D minimisation, per frame. Every meV of δ is a meV the label sits high (it describes the ideal lattice at a volume the configuration does not adopt). Prediction (4): δ > 0 on most frames with mean ≥ 5 meV/atom; falsified if the mean is < 1.

EXPERIMENTS.md · lines 15715–15742

E228 RESULT (2026-09-22 19:2x): MACE is not soft — falsified; the labeller's volume fit misses by +17 meV/atom — confirmed. 150 frames, both controls at 0.0000 meV/atom. (1) falsified: force slope s_F = 1.014 (cosine 0.987; harmonic-work ratio 0.998); by element Cr 0.94, Hf 0.96, Mo 1.05, Nb 1.06, Ta 1.03, Ti 0.98, V 1.00, W 1.01, Zr 1.01. (2) falsified: virial slope s_V = 1.072 (corr 0.978, p_DFT from −2.3 to +2.3 eV/atom); the +146 meV/atom intercept is an equilibrium-volume offset worth ≈ 0.2 % in a, ≈ 0.5 meV/atom. (3) falsified: implied softening bias −0.2 meV/atom (p10–p90 −1.1…+1.1). (4) confirmed: the stored a_relaxed is not MACE's minimum — E(ideal, a_in) < E(ideal, a_relaxed) on 128 of 150 frames; a_relaxed − a_min = +0.047 Å on average (max 0.078); δ = +17.2 meV/atom mean (median 16.8, p90 27.6, max 35.3). The parabola over ±10 % in a puts its vertex on the soft (expanded) side of a curve that is steep under compression, so every alloy label sits high by δ. What this does and does not explain. It does not explain the +32 on the 16-atom Mo–Ta cells: the table of v5's signed residuals on the 21 QE store rows (all v7b could score) shows no offset on 54-atom random- like cells — −4 ± 11 over 11 cells — while the four 16-atom Mo–Ta cells (E191; k 7³, v3 references, matching Widom's PBE B2 −186 and random −90…−127) sit at +23…+40. Those four rows were in v7b's training set; pulling the Mo–Ta region down by that much is v7b's −7.7 on the held-out Mo–Nb–Ta–W rows, without any VASP-vs-QE frame offset. The pre-registered per-group offset is therefore the wrong fix, and the "+32/+40 on random Mo–Ta cells" this record cited as a frame offset is a 16-atom effect (or a v2/v3 standard effect — E180, the 54-atom comparator, is on the older 50/400 Ry, k 3³ standard; E192_MoTa_T800, 54-atom on the v3 standard, is the control and is being scored next). Also found: the store holds rows on two DFT standards (v2: 50/400 Ry, refs k 10³; v3: 60/720 Ry, refs k 13³), each internally consistent (alloy and references on the same cutoff and k-density), whose references differ by +8.6 (Ta) to +42.2 (Nb) meV/atom in absolute energy — harmless per row, a precision difference between rows. Still open from this: δ on the pure-reference frames (the part of δ that survives into formation energies is δ_alloy − Σ xᵢ δ_pure,i), and the references' selection — each element's is the lowest corrected label among its pure frames, 3.5 (W) to 38.4 (Ti) meV/atom below its least-corrected frame, a winner's-curse estimate.

Related entries

Built with PRISMWebsite and visualizations made using Claude