Experiments · E150

Can RHEA's rattled and squeezed energies be corrected into clean training labels?

Partly. Formation energies then centred at 0.0 meV/atom instead of +622, but spread wider than predicted until strained frames were cut.

In the log: E149 found the smaller of the two defects. The labels were the bigger one.

falsifiedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 1 · ordering3 predictions · 2 result paragraphsEXPERIMENTS.md lines 9067–9132, lines 9134–9146, lines 9148–9158, lines 9752–9766
exp E150 diagram
What E150 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E150.svg).

Pre-registration

  1. (1)
    Formation energies "roughly -150 to +150": mean 0.0 but sd 259, max +6,534 — a tail of frames the correction could not bridge (max |correction| 6,580); falsified on the tails, and the same cut removes it.
    no verdict written against it
  2. (2)
    Fitted references within ~50 meV of the eight directly readable: gaps -223..+207, sd 139 — falsified, and the sign pattern is the useful part. Mo +207, W +165, Nb +98, Ta +71 sit where they should (the raw cells are rattled and lie above the fitted ideal-lattice value), but Cr -223, Hf -137, V -53 have the fitted reference above a raw rattled cell, which an elemental energy cannot do. The regression is absorbing the volume residual of diagnostic 1 along composition-correlated directions.
    no verdict written against it
  3. (3)
    Zr: no direct value to compare; -8.173 eV/atom, unverifiable until Jan's VASP or a QE bcc-Zr reference exists. Branch opened: choose the cut by measurement, not by eye. Add a --max-correction filter to the gate and the reference fit, sweep it, and take the largest cut whose residual over 0.1 A clears 15 meV/atom. Then re-fit the references on that set and re-read the gaps — the sign split should close if the leak was the tail.
    no verdict written against it
The pre-registration, as written

E150's own predictions, scored on the full set. (1) Formation energies "roughly -150 to +150": mean 0.0 but sd 259, max +6,534 — a tail of frames the correction could not bridge (max |correction| 6,580); falsified on the tails, and the same cut removes it. (2) Fitted references within ~50 meV of the eight directly readable: gaps -223..+207, sd 139 — falsified, and the sign pattern is the useful part. Mo +207, W +165, Nb +98, Ta +71 sit where they should (the raw cells are rattled and lie above the fitted ideal-lattice value), but Cr -223, Hf -137, V -53 have the fitted reference above a raw rattled cell, which an elemental energy cannot do. The regression is absorbing the volume residual of diagnostic 1 along composition-correlated directions. (3) Zr: no direct value to compare; -8.173 eV/atom, unverifiable until Jan's VASP or a QE bcc-Zr reference exists.

Branch opened: choose the cut by measurement, not by eye. Add a --max-correction filter to the gate and the reference fit, sweep it, and take the largest cut whose residual over 0.1 A clears 15 meV/atom. Then re-fit the references on that set and re-read the gaps — the sign split should close if the leak was the tail.

Results

EXPERIMENTS.md · line 9134

E150 interim, 130 of 6,449 frames. Prediction 1 confirmed: formation energies now centre at 0.0 meV/atom, sd 99, range -168 to +234, against +622 mean for the displacement-only version. That is the range a bcc solid-solution formation energy occupies.

Prediction 2 is heading for falsification and the reason may be benign. The fitted references sit 95 to 335 meV/atom below the eight read directly from RHEA's own bcc cells, against a predicted 50. Every one is lower, not scattered. The likely explanation is that the direct values were minima over the cells RHEA happens to contain, which is an upper bound on the true bcc minimum, while the fitted values come from volume-relaxed labels that reach it — in which case the two are not the same quantity and the prediction was mis-stated rather than the pipeline wrong. Not yet settled: these 130 frames are all binaries at a single compressed input lattice constant of 2.900 Å, so the regression is badly conditioned and only five of nine columns are determined. Re-score on the full set.

EXPERIMENTS.md · line 9148

E150 addendum — a claim of mine withdrawn. Reporting the published-T_c table I wrote that the ladder counts only the above-window route to a pass. That is wrong. passes_window sums below + above and integrates the scale prior over both, and tests/test_requirement.py::test_the_window_verdict_is_a_window_not_a_ranking exists precisely to stop the verdict degenerating into "hotter is better". Measured: 30 K scores 0.987, 500 K scores 0.000, 1400 K scores 1.000. Nothing needs changing.

What survives is the part that came from the literature rather than from me: 7 of the 12 published determinations have a transition inside 90-1000 K, so the requirement is hard independently of anything this project computed, and CrTaVW at 1200-1300 K is a published example of an above-window pass the ladder would score as one.


The full record

This entry is written in 4 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 9067–9132

E150 — E149 found the smaller of the two defects. The labels were the bigger one.

E149 said the eCE was training on the wrong half of RHEA and launched ece_rhea_v2_ordered to fix it. Both runs are now stopped, because an independent review found that neither training set had usable labels in the first place.

A cluster expansion is a function of the configuration and nothing else — no displacement degree of freedom, no volume degree of freedom. RHEA carries both on purpose, because it exists to train interatomic potentials. Measured over the 6,449 cubic bcc supercells:

displacement   mean 0.13 A off the ideal site       ~200 meV/atom
volume         lattice constant 2.665 to 3.818 A    up to ~1700 meV/atom
---------------------------------------------------------------------
the ordering signal the model exists to resolve       10-30 meV/atom

Neither nuisance term is a function of the configuration, so the model cannot represent them. They go into the residual. E149's controlled comparison of two training sets was a comparison of two ways of fitting noise.

This is the same error as E128 and E129, for the third time: differencing two energies whose lattice and relaxation conventions do not match. The first was on the validation side. This one was on the training side, and it was in the training set I built two days after writing the standing rule about it.

The fix is scripts/expansion/rhea_labels.py:

E_lattice = E_DFT - [ MACE(as RHEA has it) - MACE(ideal sites, relaxed volume) ]

Same cell shape, same occupation, so the subtraction removes displacement-plus-strain and nothing else. MACE appears only as a difference over one structure — 1-4 meV/atom by E78 — never as an absolute energy. The labels stay first-principles. A first version corrected only the displacement and produced formation energies averaging +622 meV/atom, which is how the volume term was found.

Two checks are carried per structure rather than assumed:

  • vertex_interior — whether the volume minimum was bracketed. RHEA samples cells up to 18% off equilibrium, so the first scan misses on about a quarter of frames and a second scan is centred on the new estimate. All 30 smoke-test frames converge after the retry.
  • harmonic_comparable / disagree_meV — RHEA ships DFT forces, giving a MACE-free estimate of the displacement part as -½ΣF·u. It applies only where the frame was already near its own equilibrium volume; elsewhere the check does not apply and is recorded as not applying, rather than being allowed to condemn the frame.

The references are fitted, not looked up. Eight of the nine elements have a bcc supercell somewhere in RHEA. Zr has none outside liquid snapshots — its lowest single-element cubic cell sits at 1.82 Å displacement — and Zr appears in 33% of the bcc frames, so dropping it is not an option. --fit-references regresses the corrected E_lattice on composition and recovers all nine under one convention.

Predicted, before the run finishes:

  1. Formation energies will land in roughly -150 to +150 meV/atom once the references are fitted, against +622 mean for the displacement-only version. That is the range a bcc solid-solution formation energy occupies and the first sign the convention is right.
  2. The fitted elemental references will sit within ~50 meV/atom of the eight that can be read directly from RHEA's own bcc cells. They are derived by completely different routes, so agreement is a real check and disagreement means the regression is absorbing something it should not.
  3. Zr's fitted reference will be the least well determined, since it never appears alone.

Falsified if the formation energies still centre far from zero, or if the eight directly readable references disagree with the fitted ones by more than ~50 meV/atom — either would mean a convention is still unmatched and the labels are not yet usable.

runs/rhea_labels.jsonl, checkpointed per structure, ~0.9 s each, 6,449 frames.

EXPERIMENTS.md · lines 9134–9146

E150 interim, 130 of 6,449 frames. Prediction 1 confirmed: formation energies now centre at 0.0 meV/atom, sd 99, range -168 to +234, against +622 mean for the displacement-only version. That is the range a bcc solid-solution formation energy occupies.

Prediction 2 is heading for falsification and the reason may be benign. The fitted references sit 95 to 335 meV/atom below the eight read directly from RHEA's own bcc cells, against a predicted 50. Every one is lower, not scattered. The likely explanation is that the direct values were minima over the cells RHEA happens to contain, which is an upper bound on the true bcc minimum, while the fitted values come from volume-relaxed labels that reach it — in which case the two are not the same quantity and the prediction was mis-stated rather than the pipeline wrong. Not yet settled: these 130 frames are all binaries at a single compressed input lattice constant of 2.900 Å, so the regression is badly conditioned and only five of nine columns are determined. Re-score on the full set.

EXPERIMENTS.md · lines 9148–9158

E150 addendum — a claim of mine withdrawn. Reporting the published-T_c table I wrote that the ladder counts only the above-window route to a pass. That is wrong. passes_window sums below + above and integrates the scale prior over both, and tests/test_requirement.py::test_the_window_verdict_is_a_window_not_a_ranking exists precisely to stop the verdict degenerating into "hotter is better". Measured: 30 K scores 0.987, 500 K scores 0.000, 1400 K scores 1.000. Nothing needs changing.

What survives is the part that came from the literature rather than from me: 7 of the 12 published determinations have a transition inside 90-1000 K, so the requirement is hard independently of anything this project computed, and CrTaVW at 1200-1300 K is a published example of an above-window pass the ladder would score as one.

EXPERIMENTS.md · lines 9752–9766

E150's own predictions, scored on the full set. (1) Formation energies "roughly -150 to +150": mean 0.0 but sd 259, max +6,534 — a tail of frames the correction could not bridge (max |correction| 6,580); falsified on the tails, and the same cut removes it. (2) Fitted references within ~50 meV of the eight directly readable: gaps -223..+207, sd 139 — falsified, and the sign pattern is the useful part. Mo +207, W +165, Nb +98, Ta +71 sit where they should (the raw cells are rattled and lie above the fitted ideal-lattice value), but Cr -223, Hf -137, V -53 have the fitted reference above a raw rattled cell, which an elemental energy cannot do. The regression is absorbing the volume residual of diagnostic 1 along composition-correlated directions. (3) Zr: no direct value to compare; -8.173 eV/atom, unverifiable until Jan's VASP or a QE bcc-Zr reference exists.

Branch opened: choose the cut by measurement, not by eye. Add a --max-correction filter to the gate and the reference fit, sweep it, and take the largest cut whose residual over 0.1 A clears 15 meV/atom. Then re-fit the references on that set and re-read the gaps — the sign split should close if the leak was the tail.

Related entries

Built with PRISMWebsite and visualizations made using Claude