Experiments · E151

Are the corrected labels real signal, or mostly the correction's own error?

Partly. Below 0.10 Å of strain they are usable: noise 12.4 meV/atom against a 29 meV/atom ordering signal.

In the log: Is the corrected label a label, or is it the correction's error?

mixedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 4 · DFT3 predictions · 7 result paragraphsEXPERIMENTS.md lines 9162–9209, lines 9715–9750, lines 9768–9796, lines 9798–9824, lines 9826–9859, lines 9861–9905, lines 9907–9929, lines 9931–9964, lines 9966–10011, lines 10584–10597
exp E151 diagram
What E151 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E151.svg).

Pre-registration

  1. (1)
    its within-composition rank correlation with DFT is below 0.3;
    no verdict written against it
  2. (2)
    its |dE| scale is under a third of DFT's (DFT's is ~29 meV/atom here) — the "too flat" reading of E150's interim, now on labels that can support it;
    no verdict written against it
  3. (3)
    had bcc cells inside bcc_binary_alloys: Hf: MACE minimum 3.527 A vs DFT's lowest raw cell at 3.600, -0.073 A — outside 0.03, but Hf was the element I said would be undefined, since its bcc "minimum" is the instability; Ti: -0.005 A, inside 0.01. The test's actual target, Cr, is pending a label pass over pure_elem_lowE (launched now, ~105 cubic cells). The P1 training file. runs/rhea_labels_formation_strain0.10.jsonl: 4,326 frames, formation energies mean 0.0, sd 63.4, range -305 to +189 meV/atom, references Cr -9.243, Hf -12.406, Mo -11.045, Nb -10.185, Ta -11.766, Ti -7.703, V -8.895, W -12.950, Zr -8.183 eV/atom (within 5 meV of the cut-400/800 values, as the non-tail mechanism requires). P2 baseline, predicted before it runs. Scoring ce_12element — the current bottom rung, fitted to MACE labels on random cells, never shown an ordered cell — on these labels, within composition where every offset cancels
    no verdict written against it
The pre-registration, as written

E151, three closures.

The eleven non-finite frames are the +6.5 eV tail. Three 128-atom ordered cells and eight 54-atom random ones have a_relaxed = NaN and corrections of -4.2 to -6.6 eV/atom — MACE returned non-finite energies somewhere in their volume scans (input lattice constants 3.06 to 3.76 A, the expanded end), the parabola fit propagated it, and E_lattice = E_DFT - (-6 eV) became the +6,534 meV/atom maximum E150's prediction 1 was falsified on. So that tail was eleven broken frames, not a strained population; the strain cut and an isfinite guard in the reference pass both exclude them. E150 prediction 1 stands falsified on the sd (259) but its "max +6,534" was a bug, not a distribution.

The Cr-volume test cannot run as pre-registered. The label pass covered the three alloy groups; the pure-element cubic cells for Cr, Mo, Nb, Ta, V and W sit in pure_elem_lowE, which it skipped. Only Hf (12 frames) and Ti (3) had bcc cells inside bcc_binary_alloys: Hf: MACE minimum 3.527 A vs DFT's lowest raw cell at 3.600, -0.073 A — outside 0.03, but Hf was the element I said would be undefined, since its bcc "minimum" is the instability; Ti: -0.005 A, inside 0.01. The test's actual target, Cr, is pending a label pass over pure_elem_lowE (launched now, ~105 cubic cells).

The P1 training file. runs/rhea_labels_formation_strain0.10.jsonl: 4,326 frames, formation energies mean 0.0, sd 63.4, range -305 to +189 meV/atom, references Cr -9.243, Hf -12.406, Mo -11.045, Nb -10.185, Ta -11.766, Ti -7.703, V -8.895, W -12.950, Zr -8.183 eV/atom (within 5 meV of the cut-400/800 values, as the non-tail mechanism requires).

P2 baseline, predicted before it runs. Scoring ce_12element — the current bottom rung, fitted to MACE labels on random cells, never shown an ordered cell — on these labels, within composition where every offset cancels: (1) its within-composition rank correlation with DFT is below 0.3; (2) its |dE| scale is under a third of DFT's (DFT's is ~29 meV/atom here) — the "too flat" reading of E150's interim, now on labels that can support it; (3) MACE-MPA-0 on the same structures ranks better than the expansion (rho above 0.5), since it at least saw ordered intermetallics in training. Falsified if the expansion's rho/comp exceeds 0.6, which would mean it resolves ordering it was never trained on and the case for replacing it weakens to a cost argument.

E151 follow-up — references from corrected pure cells, done. --pure-references holds Cr, Mo, Nb, Ta, V, W at the lowest corrected pure bcc cell from the pure_elem_lowE pass and regresses only Hf, Ti, Zr on the remainder:

Cr -9.5084   Mo -10.9154   Nb -10.2122   Ta -11.7940   V -8.9864   W -12.9257   (measured)
Hf -12.3721  Ti -7.6645    Zr -8.1505                                            (fitted)

formation energy, 4,326 frames at the 0.10 A cut:  mean +17.2  sd 83.1  min -270  max +324 meV/atom

The mean is no longer zero by construction — a set of real bcc formation energies has a mean, and this one is positive because the equation-of-state sets sample unfavourable mixtures on purpose. This file, rhea_labels_formation_strain0.10_pureref.jsonl, is the one a hull is built from. Ordering energies are unchanged by a per-element shift; the eCE training file need not be rebuilt for P1, only for any formation-energy use.


Results

EXPERIMENTS.md · line 9715

E151 result — the label gate, on all 6,449 corrected frames. Prediction 1 falsified; the labels are not usable as-is, and the reason is the strained tail, not the correction.

|correction|: median 61, p90 726, max 6,580 meV/atom

1. residual volume dependence, within composition (100 compositions, >=4 frames, >0.05 A)
     |dE/da| before   1,962 meV/atom per A
     |dE/da| after      314 meV/atom per A      84.0% removed
     over 0.1 A          31.4 meV/atom          gate was 15  ->  NOT usable as-is

3. within-composition spread against correction size
     |corr| < 100    3,806 frames   mean |dE|  29.3
     |corr| < 200    4,603          28.8
     |corr| < 400    5,277          28.1
     |corr| < 800    5,878          31.4
     all             6,449          44.4

Prediction 1 falsified: 84% removed and 31.4 meV/atom residual, against ">90%" and "under 15". Not the abandon-the-correction branch either — I set that at a residual above 30% of the uncorrected slope, and it is 16%.

Prediction 2 resolved toward physics, which is the outcome I said would be the more interesting one. The spread sits at 28-31 meV/atom for every cut up to 800 and only rises to 44 when the 571 heaviest-corrected frames are admitted. A spread that does not move across a factor-of-eight range of correction size is not correction error. It is the energy difference between ordered and random configurations at one composition — the bcc_alloys_ordered 16-atom cells against the 54-atom solid solutions sharing a composition key — and ~29 meV/atom is the textbook scale of a bcc refractory ordering energy. This is the quantity the ladder's bottom rung exists to resolve, measured directly from DFT for the first time in this project. Whether the current expansion reproduces it is the P2 scoring on this set, not yet run on the full labels; the interim "x45 too flat" of E150 was measured on random-only binaries where the true spread is small and the correction error was not, and is withdrawn as a comparison.

Prediction 3 confirmed: a strain cut is required, and the tail it removes is where bcc_binary_alloys samples the equation of state on purpose.

EXPERIMENTS.md · line 9768

E151, the cut sweep — and the gate quantity fails as a cut-selector. Diagnostic 1 at increasing strain cuts (|correction| in meV/atom, --skip-direct):

cut     kept    removed    residual over 0.1 A
100    3,806    44.3%      75.3
200    4,603    71.1%      38.1
400    5,277    63.8%      17.7
800    5,878    69.2%      17.9
none   6,449    84.0%      31.4

Non-monotonic, and the tightest cut is the worst. That is not the labels; it is the metric. Diagnostic 1 fits a linear slope of E_lattice against a_in within a composition. Under a tight cut the surviving frames are the ones already near their own equilibrium, where E(a) is quadratic and nearly flat and the span of a_in is narrow, so a straight line through them measures curvature and fitting noise, not a residual volume dependence — and the "removed %" shrinks because the before slope is small there too, not because the correction got worse. The pre-registered gate quantity is a sound test for "did the correction do anything" on the full set, and an unsound one for choosing where to cut. I stated it as the decisive diagnostic; as a selector it is withdrawn.

What is robust in the table: 400 and 800 sit on a plateau at ~18, and the jump to 31 is the

800 tail — the same 571 frames diagnostic 3 identifies as where correction error enters.

Replacement, stated before it is run: within composition, regress E_lattice on the signed input strain d = a_in - a_relaxed and report slope x median|d| — the label's remaining dependence on how strained the input frame was, which is what the correction was supposed to remove and which does not conflate with the physical curvature of E(a). Predicted: monotone in the cut, under 15 meV/atom by cut 800, and the reference-gap sign split (Cr, Hf, V fitted above raw cells) closes at the same cut.

EXPERIMENTS.md · line 9798

E151, diagnostic 1b — my replacement is falsified by the same mechanism, and the mechanism is now identified. Residual over a typical frame, regressing the label on input strain a_in - a_relaxed within composition:

cut    kept    residual (meV/atom)
100   3,806    78.9
200   4,603    66.3
400   5,277    30.0
800   5,878    36.7

I predicted monotone and under 15 by cut 800. It is neither. Tighter cuts are worse again, and the arithmetic says why: at cut 100 a residual of 79 over a median strain of ~0.03 A means a fitted slope near 2,600 meV/atom per A, which no correction error produces. What produces it is the ~29 meV/atom ordered-versus-random spread measured in diagnostic 3 — physics, uncorrelated with strain — divided by a strain span of a few hundredths of an angstrom. Whenever 16-atom ordered cells share a composition key with 54-atom random cells, any within-composition regression on a narrow span reads the configuration effect as a strain slope. Diagnostic 1 had the same contamination and hid it under the curvature story.

The separation the diagnostics need: restrict the strain test to the random groups (bcc_alloys, bcc_binary_alloys). At one composition, random 54-atom decorations have nearly identical cluster vectors and therefore nearly identical energies — that is the literature result E149 rests on — so their residual variation is strain or correction error, with no configuration term to confuse it. The ordered cells are also the lightly corrected ones (Fable measured -46 +/- 12 meV/atom of rattle on them), so the strain question was always about the equation-of-state sets. Predicted, before running: on random-only groups the residual is monotone in the cut and under 15 meV/atom by cut 400.

EXPERIMENTS.md · line 9826

E151, diagnostic 1b on random groups only — falsified a third time, and the reason is structural.

cut    compositions testable    residual (meV/atom)
none          78                 43.3
100            0                 (nothing left to test)
200            2                 25.1
400           32                 20.9
800           61                 32.4

Restricting to random groups removes the absurd slopes — 79 becomes 21 at the best cut — so the mechanism named above was right. But "monotone and under 15 by 400" is wrong: 400 is a minimum at 20.9, and tightening past it does not lower the residual, it removes the test set, because the strained equation-of-state frames are both what tests the correction and what the cut discards. A cut on |correction| cannot be selected by this diagnostic, on principle, not just on these numbers.

What is measured and stands: at the plateau, the correction leaves ~21 meV/atom of strain-correlated error on a typical random EOS frame — MACE's relative accuracy over a few-hundred-meV bridge is 5-8%, and 5-8% of that bridge is the size of the ordering signal. By the same slope an ordered cell at its typical ~0.03 A of input strain carries ~6 meV/atom. The contamination is concentrated in the frames that carry the least ordering information.

Reframed, with a prediction before the measurement. Cut on the input strain |a_in - a_relaxed| itself, not on the correction, and judge the cut by the cleanest noise estimate available: the within-composition sd among random decorations, which by the cluster-vector argument have almost no physical spread, so their post-correction spread is correction error and nothing else. Predicted: that sd is under 10 meV/atom at |strain| <= 0.05 A and rises monotonically with the cut; at least 80% of the bcc_alloys_ordered cells survive a 0.05 A cut (they were rattled, not strained); and the random EOS frames beyond 0.10 A are where the sd exceeds 20. Falsified if the sd among random decorations is already above 15 at 0.05 A — then the residual is not strain-correlated at all and the correction is noisy even where it has little to do, which would send this back to the label pipeline rather than to a cut.

EXPERIMENTS.md · line 9861

E151, diagnostic 2 under the cut — the sign split does not close, and my prediction that it would is falsified. Reference gaps (fitted minus raw direct, meV/atom):

el     none    cut 400   cut 800
Cr     -223    -200      -195
Hf     -137    -106       -97
Mo     +207    +194      +193
Nb      +98     +28       +40
Ta      +71     +22       +27
Ti       +4     +12        +6
V       -53     -52       -42
W       +165    +116      +129
sd       139     115       115

Nb, Ta, W and Ti tighten with the cut, as a strained-tail leak would. Cr, Mo, Hf and V do not move. So the tail was part of it and the rest is something a strain cut cannot reach, and the elements name it:

  • Hf, and to a lesser extent Ti and Zr: physics. Group-IV bcc is dynamically unstable at 0 K. A rattled bcc Hf cell is not a noisy version of the ideal one, it is partway down a real instability, so its raw energy sits below the ideal-lattice reference. A fitted reference above a raw cell is exactly what should happen for Hf. This is the blind spot already recorded against every bcc cluster expansion in this project, appearing in the references.
  • Cr, and probably V: the convention sits at MACE's volume, not DFT's. E_lattice is the energy of the ideal configuration at the volume MACE relaxes it to. Where MACE's equation-of-state minimum is displaced from PBE's, every frame containing that element carries an offset proportional to its fraction — a composition-dependent shift, which is precisely the linear-in-composition term the regression puts into the reference. RHEA is non-spin-polarised and MACE-MPA-0 was trained on a Materials Project chromium that is not, so Cr is the element where the two equilibria are least likely to agree.
  • Mo +194 and W +193 persist too and are not group IV; they are the largest positive gaps, consistent with the same volume mismatch in the other direction.

What this does and does not cost. A per-element volume offset is absorbed exactly by a cluster expansion's singlet terms and cancels in any energy difference at fixed composition. It does not touch ordering energies, which is what the ladder's bottom rung is for. It does shift formation energies and therefore the hull, so the hull must not be built from these labels without the offset measured.

Pre-registered test of the Cr mechanism. For pure-element frames the label file records a_in (RHEA's cell) and a_relaxed (MACE's minimum). Take, per element, the a_in of the lowest-energy raw cubic cell as the PBE equilibrium and compare it with MACE's a_relaxed for the same frames. Predicted: Cr differs by more than 0.03 A; Mo, Nb, Ta, W within 0.01 A; Hf and Ti undefined in the sense that their bcc minimum is not a minimum. If Cr agrees with DFT to 0.01 A the mechanism is wrong and the Cr offset is unexplained.

EXPERIMENTS.md · line 9907

E151 result — the cut, chosen by the noise among random decorations.

|strain| <=   kept   random  ordered (kept%)   testable comps   sd among random decorations   max |corr|
0.02 A         399     215      184   ( 7%)          1                 0.4                        170
0.05 A       2,068     835    1,233   (46%)          1                 0.4                        198
0.10 A       4,326   1,708    2,618   (97%)         14                12.4                        327
0.20 A       5,399   2,711    2,688  (100%)        132                47.1                      1,162
all          6,438   3,750    2,688  (100%)        703                54.2                      4,537

Monotone: confirmed. The spread among random decorations of one composition — which have almost no physical spread, so this is correction error — rises 0.4 → 12.4 → 47 → 54 with the strain admitted. "At least 80% of ordered cells survive 0.05 A": falsified, 46%. The ordered cells are more strained relative to MACE's minimum than "rattled, not strained" implied. "Under 10 at 0.05 A": unassessable, one testable composition.

The cut is 0.10 A. It keeps 4,326 frames including 97% of the ordered cells, caps the correction at 327 meV/atom, and leaves a label noise floor of 12.4 meV/atom against the 29 meV/atom ordering spread of diagnostic 3 — about 2:1 per structure, which a fit over thousands of structures averages down and which is stated here so nobody reads the training loss as the label accuracy. Past 0.10 A the error is ~50 meV/atom and would swamp the signal; the frames it removes are the equation-of-state sets, not the ordered ones.

Eleven frames fall out of "all" (6,438 of 6,449): their strain is non-finite. Checked below.

EXPERIMENTS.md · line 9966

E151, P2 baseline — my prediction is falsified on every count, and the falsification is the most important number of the day. ce_12element and MACE-MPA-0 scored within composition on the 0.10 A labels (212 compositions with >= 3 configurations, 2,292 structures):

model           MAE dE   rho pooled   rho/comp   |dE| DFT   |dE| model
ce_12element     14.5      0.847        0.651      28.8       28.1
MACE-MPA-0       12.0      0.895        0.671      28.8       28.9
label noise floor (E151)                                       12.4

(1) "rho/comp below 0.3": 0.651. (2) "|dE| under a third of DFT's": 28.1 against 28.8 — the scale is right to 2%. (3) "MACE ranks better": true by 0.02, which is nothing. The falsification condition I set — rho/comp above 0.6 — is met. The current bottom rung resolves ordering energies it was never shown, at the right magnitude, and its 14.5 meV/atom error is at the 12.4 meV/atom noise floor of the labels themselves. On this set it cannot be told apart from a perfect model, and neither can MACE.

Withdrawn: E150's interim "the expansion is 45x too flat" and E151's "17x" — both were measured on random-only binaries where the DFT spread was correction error, not physics. The expansion's spread matches DFT's once the labels are clean. The case for the eCE is not within-distribution ordering accuracy; it is extrapolation into withheld chemistry and cost per element, which is what the paper claims and what P2's MoNbTaW/MoNbTaVW holdout tests.

Cr-volume test — falsified, informatively. MACE's bcc minimum sits above DFT's lowest raw cell for all six elements, not Cr alone: Cr +0.057, Mo +0.035, Nb +0.030, Ta +0.050, V +0.041, W +0.037 A. A uniform ~1% offset is worth tens of meV and has one sign; it cannot make Cr -200 against Mo +190.

The two-route reference comparison, redone in one convention — fitted references at the 0.10 A cut against the lowest corrected pure cell from the pure_elem_lowE pass:

Cr +265   Mo -129   V +91   Nb +27   Ta +28   W -25   meV/atom (fitted minus pure)

Still scattered, and pure Cr cells needed only 63 meV/atom of correction, so this is not the correction either. The mechanism, and the one I recorded in memory and on the plan page is withdrawn: a composition-linear regression over alloys does not recover elemental energies. Its intercepts are "elemental energy plus the mean mixing slope along that element's direction", and with five pure-Cr frames against thousands of Cr-containing alloys the Cr column is set by the alloys. That is a property of regression intercepts, not of MACE's volume. The right references are the corrected pure cells where they exist (six elements) with only Hf, Ti and Zr fitted — a follow-up for the hull, since a per-element constant is absorbed exactly by any expansion's point terms and does not touch ordering.

P1 died on my bug. train_ece.py wrote each m x m x m supercell's fractional coordinates into a 1 x 1 x 1 cell — 54 atoms in a 3.29 A cube — and pyeCE rejected all 4,326 structures as "lattice deformation too large". ideal_frac_of returns m; emit discarded it. Fixed; relaunched to a fresh directory.


The full record

This entry is written in 10 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 9162–9209

E151 — Is the corrected label a label, or is it the correction's error?

E150 removes displacement and volume energy with one MACE difference per structure. On the first 529 frames that correction has a median magnitude of 657 meV/atom and reaches 2,941. The ordering signal is 10-30 meV/atom. So the whole pipeline rests on a quantity nobody has measured: MACE's relative accuracy on a correction of that size. At 1% it is fine. At 5% the labels carry 33 meV/atom of noise and the exercise is pointless.

E78's 1-4 meV/atom was measured on small differences. It does not transfer to these.

The first symptom is already visible. Scoring ce_12element on the partial labels, the within-composition energy spread is 76.3 meV/atom in the corrected DFT and 1.7 meV/atom in the expansion — a factor of 45. Either the expansion is far too flat, or most of that 76 is correction error, or both. The scorer cannot tell those apart and neither can a validation curve.

The test that can. A correct volume correction has one signature that cannot be faked: after it, the label must not depend on the input lattice constant. RHEA samples the same composition at many lattice constants, so within a composition the residual slope dE_lattice/da_in is a direct read of what the correction failed to remove. Three diagnostics, run together as the gate on E150:

  1. residual |dE/da| within composition, against the same slope before correction;
  2. the eight elemental references that can be read directly from RHEA's own bcc cells, against the nine fitted by regression;
  3. the within-composition spread as a function of correction size — if the spread is correction error it shrinks when restricted to small-correction frames, and if it is physics it does not.

Predicted:

  1. The correction removes more than 90% of the volume dependence, leaving under ~15 meV/atom across a typical 0.1 Å spread. Below that the labels are usable; above it they are not, whatever the training loss says.
  2. The within-composition spread falls with correction size — the frames needing under 200 meV/atom of correction will show a spread well below the pooled 76 meV/atom. If it does not fall, the 76 is physics and the expansion really is 45 times too flat, which is a far more interesting result and would be the strongest argument yet for replacing it.
  3. A strain cut will be necessary, keeping frames whose input lattice constant is within ~0.05 Å of relaxed, and it will cost a large fraction of bcc_binary_alloys, which is an equation-of-state set sampled far from equilibrium on purpose.

Falsified if the residual slope stays above ~30% of the uncorrected one, which would mean MACE cannot bridge these volumes and the honest move is to abandon the correction and train on near-equilibrium frames only — fewer structures, but labels that mean something.

Deliberately not assumed: prediction 2 has two outcomes and both are informative, so this is not a check being run to confirm something already believed.

EXPERIMENTS.md · lines 9715–9750

E151 result — the label gate, on all 6,449 corrected frames. Prediction 1 falsified; the labels are not usable as-is, and the reason is the strained tail, not the correction.

|correction|: median 61, p90 726, max 6,580 meV/atom

1. residual volume dependence, within composition (100 compositions, >=4 frames, >0.05 A)
     |dE/da| before   1,962 meV/atom per A
     |dE/da| after      314 meV/atom per A      84.0% removed
     over 0.1 A          31.4 meV/atom          gate was 15  ->  NOT usable as-is

3. within-composition spread against correction size
     |corr| < 100    3,806 frames   mean |dE|  29.3
     |corr| < 200    4,603          28.8
     |corr| < 400    5,277          28.1
     |corr| < 800    5,878          31.4
     all             6,449          44.4

Prediction 1 falsified: 84% removed and 31.4 meV/atom residual, against ">90%" and "under 15". Not the abandon-the-correction branch either — I set that at a residual above 30% of the uncorrected slope, and it is 16%.

Prediction 2 resolved toward physics, which is the outcome I said would be the more interesting one. The spread sits at 28-31 meV/atom for every cut up to 800 and only rises to 44 when the 571 heaviest-corrected frames are admitted. A spread that does not move across a factor-of-eight range of correction size is not correction error. It is the energy difference between ordered and random configurations at one composition — the bcc_alloys_ordered 16-atom cells against the 54-atom solid solutions sharing a composition key — and ~29 meV/atom is the textbook scale of a bcc refractory ordering energy. This is the quantity the ladder's bottom rung exists to resolve, measured directly from DFT for the first time in this project. Whether the current expansion reproduces it is the P2 scoring on this set, not yet run on the full labels; the interim "x45 too flat" of E150 was measured on random-only binaries where the true spread is small and the correction error was not, and is withdrawn as a comparison.

Prediction 3 confirmed: a strain cut is required, and the tail it removes is where bcc_binary_alloys samples the equation of state on purpose.

EXPERIMENTS.md · lines 9768–9796

E151, the cut sweep — and the gate quantity fails as a cut-selector. Diagnostic 1 at increasing strain cuts (|correction| in meV/atom, --skip-direct):

cut     kept    removed    residual over 0.1 A
100    3,806    44.3%      75.3
200    4,603    71.1%      38.1
400    5,277    63.8%      17.7
800    5,878    69.2%      17.9
none   6,449    84.0%      31.4

Non-monotonic, and the tightest cut is the worst. That is not the labels; it is the metric. Diagnostic 1 fits a linear slope of E_lattice against a_in within a composition. Under a tight cut the surviving frames are the ones already near their own equilibrium, where E(a) is quadratic and nearly flat and the span of a_in is narrow, so a straight line through them measures curvature and fitting noise, not a residual volume dependence — and the "removed %" shrinks because the before slope is small there too, not because the correction got worse. The pre-registered gate quantity is a sound test for "did the correction do anything" on the full set, and an unsound one for choosing where to cut. I stated it as the decisive diagnostic; as a selector it is withdrawn.

What is robust in the table: 400 and 800 sit on a plateau at ~18, and the jump to 31 is the

800 tail — the same 571 frames diagnostic 3 identifies as where correction error enters.

Replacement, stated before it is run: within composition, regress E_lattice on the signed input strain d = a_in - a_relaxed and report slope x median|d| — the label's remaining dependence on how strained the input frame was, which is what the correction was supposed to remove and which does not conflate with the physical curvature of E(a). Predicted: monotone in the cut, under 15 meV/atom by cut 800, and the reference-gap sign split (Cr, Hf, V fitted above raw cells) closes at the same cut.

EXPERIMENTS.md · lines 9798–9824

E151, diagnostic 1b — my replacement is falsified by the same mechanism, and the mechanism is now identified. Residual over a typical frame, regressing the label on input strain a_in - a_relaxed within composition:

cut    kept    residual (meV/atom)
100   3,806    78.9
200   4,603    66.3
400   5,277    30.0
800   5,878    36.7

I predicted monotone and under 15 by cut 800. It is neither. Tighter cuts are worse again, and the arithmetic says why: at cut 100 a residual of 79 over a median strain of ~0.03 A means a fitted slope near 2,600 meV/atom per A, which no correction error produces. What produces it is the ~29 meV/atom ordered-versus-random spread measured in diagnostic 3 — physics, uncorrelated with strain — divided by a strain span of a few hundredths of an angstrom. Whenever 16-atom ordered cells share a composition key with 54-atom random cells, any within-composition regression on a narrow span reads the configuration effect as a strain slope. Diagnostic 1 had the same contamination and hid it under the curvature story.

The separation the diagnostics need: restrict the strain test to the random groups (bcc_alloys, bcc_binary_alloys). At one composition, random 54-atom decorations have nearly identical cluster vectors and therefore nearly identical energies — that is the literature result E149 rests on — so their residual variation is strain or correction error, with no configuration term to confuse it. The ordered cells are also the lightly corrected ones (Fable measured -46 +/- 12 meV/atom of rattle on them), so the strain question was always about the equation-of-state sets. Predicted, before running: on random-only groups the residual is monotone in the cut and under 15 meV/atom by cut 400.

EXPERIMENTS.md · lines 9826–9859

E151, diagnostic 1b on random groups only — falsified a third time, and the reason is structural.

cut    compositions testable    residual (meV/atom)
none          78                 43.3
100            0                 (nothing left to test)
200            2                 25.1
400           32                 20.9
800           61                 32.4

Restricting to random groups removes the absurd slopes — 79 becomes 21 at the best cut — so the mechanism named above was right. But "monotone and under 15 by 400" is wrong: 400 is a minimum at 20.9, and tightening past it does not lower the residual, it removes the test set, because the strained equation-of-state frames are both what tests the correction and what the cut discards. A cut on |correction| cannot be selected by this diagnostic, on principle, not just on these numbers.

What is measured and stands: at the plateau, the correction leaves ~21 meV/atom of strain-correlated error on a typical random EOS frame — MACE's relative accuracy over a few-hundred-meV bridge is 5-8%, and 5-8% of that bridge is the size of the ordering signal. By the same slope an ordered cell at its typical ~0.03 A of input strain carries ~6 meV/atom. The contamination is concentrated in the frames that carry the least ordering information.

Reframed, with a prediction before the measurement. Cut on the input strain |a_in - a_relaxed| itself, not on the correction, and judge the cut by the cleanest noise estimate available: the within-composition sd among random decorations, which by the cluster-vector argument have almost no physical spread, so their post-correction spread is correction error and nothing else. Predicted: that sd is under 10 meV/atom at |strain| <= 0.05 A and rises monotonically with the cut; at least 80% of the bcc_alloys_ordered cells survive a 0.05 A cut (they were rattled, not strained); and the random EOS frames beyond 0.10 A are where the sd exceeds 20. Falsified if the sd among random decorations is already above 15 at 0.05 A — then the residual is not strain-correlated at all and the correction is noisy even where it has little to do, which would send this back to the label pipeline rather than to a cut.

EXPERIMENTS.md · lines 9861–9905

E151, diagnostic 2 under the cut — the sign split does not close, and my prediction that it would is falsified. Reference gaps (fitted minus raw direct, meV/atom):

el     none    cut 400   cut 800
Cr     -223    -200      -195
Hf     -137    -106       -97
Mo     +207    +194      +193
Nb      +98     +28       +40
Ta      +71     +22       +27
Ti       +4     +12        +6
V       -53     -52       -42
W       +165    +116      +129
sd       139     115       115

Nb, Ta, W and Ti tighten with the cut, as a strained-tail leak would. Cr, Mo, Hf and V do not move. So the tail was part of it and the rest is something a strain cut cannot reach, and the elements name it:

  • Hf, and to a lesser extent Ti and Zr: physics. Group-IV bcc is dynamically unstable at 0 K. A rattled bcc Hf cell is not a noisy version of the ideal one, it is partway down a real instability, so its raw energy sits below the ideal-lattice reference. A fitted reference above a raw cell is exactly what should happen for Hf. This is the blind spot already recorded against every bcc cluster expansion in this project, appearing in the references.
  • Cr, and probably V: the convention sits at MACE's volume, not DFT's. E_lattice is the energy of the ideal configuration at the volume MACE relaxes it to. Where MACE's equation-of-state minimum is displaced from PBE's, every frame containing that element carries an offset proportional to its fraction — a composition-dependent shift, which is precisely the linear-in-composition term the regression puts into the reference. RHEA is non-spin-polarised and MACE-MPA-0 was trained on a Materials Project chromium that is not, so Cr is the element where the two equilibria are least likely to agree.
  • Mo +194 and W +193 persist too and are not group IV; they are the largest positive gaps, consistent with the same volume mismatch in the other direction.

What this does and does not cost. A per-element volume offset is absorbed exactly by a cluster expansion's singlet terms and cancels in any energy difference at fixed composition. It does not touch ordering energies, which is what the ladder's bottom rung is for. It does shift formation energies and therefore the hull, so the hull must not be built from these labels without the offset measured.

Pre-registered test of the Cr mechanism. For pure-element frames the label file records a_in (RHEA's cell) and a_relaxed (MACE's minimum). Take, per element, the a_in of the lowest-energy raw cubic cell as the PBE equilibrium and compare it with MACE's a_relaxed for the same frames. Predicted: Cr differs by more than 0.03 A; Mo, Nb, Ta, W within 0.01 A; Hf and Ti undefined in the sense that their bcc minimum is not a minimum. If Cr agrees with DFT to 0.01 A the mechanism is wrong and the Cr offset is unexplained.

EXPERIMENTS.md · lines 9907–9929

E151 result — the cut, chosen by the noise among random decorations.

|strain| <=   kept   random  ordered (kept%)   testable comps   sd among random decorations   max |corr|
0.02 A         399     215      184   ( 7%)          1                 0.4                        170
0.05 A       2,068     835    1,233   (46%)          1                 0.4                        198
0.10 A       4,326   1,708    2,618   (97%)         14                12.4                        327
0.20 A       5,399   2,711    2,688  (100%)        132                47.1                      1,162
all          6,438   3,750    2,688  (100%)        703                54.2                      4,537

Monotone: confirmed. The spread among random decorations of one composition — which have almost no physical spread, so this is correction error — rises 0.4 → 12.4 → 47 → 54 with the strain admitted. "At least 80% of ordered cells survive 0.05 A": falsified, 46%. The ordered cells are more strained relative to MACE's minimum than "rattled, not strained" implied. "Under 10 at 0.05 A": unassessable, one testable composition.

The cut is 0.10 A. It keeps 4,326 frames including 97% of the ordered cells, caps the correction at 327 meV/atom, and leaves a label noise floor of 12.4 meV/atom against the 29 meV/atom ordering spread of diagnostic 3 — about 2:1 per structure, which a fit over thousands of structures averages down and which is stated here so nobody reads the training loss as the label accuracy. Past 0.10 A the error is ~50 meV/atom and would swamp the signal; the frames it removes are the equation-of-state sets, not the ordered ones.

Eleven frames fall out of "all" (6,438 of 6,449): their strain is non-finite. Checked below.

EXPERIMENTS.md · lines 9931–9964

E151, three closures.

The eleven non-finite frames are the +6.5 eV tail. Three 128-atom ordered cells and eight 54-atom random ones have a_relaxed = NaN and corrections of -4.2 to -6.6 eV/atom — MACE returned non-finite energies somewhere in their volume scans (input lattice constants 3.06 to 3.76 A, the expanded end), the parabola fit propagated it, and E_lattice = E_DFT - (-6 eV) became the +6,534 meV/atom maximum E150's prediction 1 was falsified on. So that tail was eleven broken frames, not a strained population; the strain cut and an isfinite guard in the reference pass both exclude them. E150 prediction 1 stands falsified on the sd (259) but its "max +6,534" was a bug, not a distribution.

The Cr-volume test cannot run as pre-registered. The label pass covered the three alloy groups; the pure-element cubic cells for Cr, Mo, Nb, Ta, V and W sit in pure_elem_lowE, which it skipped. Only Hf (12 frames) and Ti (3) had bcc cells inside bcc_binary_alloys: Hf: MACE minimum 3.527 A vs DFT's lowest raw cell at 3.600, -0.073 A — outside 0.03, but Hf was the element I said would be undefined, since its bcc "minimum" is the instability; Ti: -0.005 A, inside 0.01. The test's actual target, Cr, is pending a label pass over pure_elem_lowE (launched now, ~105 cubic cells).

The P1 training file. runs/rhea_labels_formation_strain0.10.jsonl: 4,326 frames, formation energies mean 0.0, sd 63.4, range -305 to +189 meV/atom, references Cr -9.243, Hf -12.406, Mo -11.045, Nb -10.185, Ta -11.766, Ti -7.703, V -8.895, W -12.950, Zr -8.183 eV/atom (within 5 meV of the cut-400/800 values, as the non-tail mechanism requires).

P2 baseline, predicted before it runs. Scoring ce_12element — the current bottom rung, fitted to MACE labels on random cells, never shown an ordered cell — on these labels, within composition where every offset cancels: (1) its within-composition rank correlation with DFT is below 0.3; (2) its |dE| scale is under a third of DFT's (DFT's is ~29 meV/atom here) — the "too flat" reading of E150's interim, now on labels that can support it; (3) MACE-MPA-0 on the same structures ranks better than the expansion (rho above 0.5), since it at least saw ordered intermetallics in training. Falsified if the expansion's rho/comp exceeds 0.6, which would mean it resolves ordering it was never trained on and the case for replacing it weakens to a cost argument.

EXPERIMENTS.md · lines 9966–10011

E151, P2 baseline — my prediction is falsified on every count, and the falsification is the most important number of the day. ce_12element and MACE-MPA-0 scored within composition on the 0.10 A labels (212 compositions with >= 3 configurations, 2,292 structures):

model           MAE dE   rho pooled   rho/comp   |dE| DFT   |dE| model
ce_12element     14.5      0.847        0.651      28.8       28.1
MACE-MPA-0       12.0      0.895        0.671      28.8       28.9
label noise floor (E151)                                       12.4

(1) "rho/comp below 0.3": 0.651. (2) "|dE| under a third of DFT's": 28.1 against 28.8 — the scale is right to 2%. (3) "MACE ranks better": true by 0.02, which is nothing. The falsification condition I set — rho/comp above 0.6 — is met. The current bottom rung resolves ordering energies it was never shown, at the right magnitude, and its 14.5 meV/atom error is at the 12.4 meV/atom noise floor of the labels themselves. On this set it cannot be told apart from a perfect model, and neither can MACE.

Withdrawn: E150's interim "the expansion is 45x too flat" and E151's "17x" — both were measured on random-only binaries where the DFT spread was correction error, not physics. The expansion's spread matches DFT's once the labels are clean. The case for the eCE is not within-distribution ordering accuracy; it is extrapolation into withheld chemistry and cost per element, which is what the paper claims and what P2's MoNbTaW/MoNbTaVW holdout tests.

Cr-volume test — falsified, informatively. MACE's bcc minimum sits above DFT's lowest raw cell for all six elements, not Cr alone: Cr +0.057, Mo +0.035, Nb +0.030, Ta +0.050, V +0.041, W +0.037 A. A uniform ~1% offset is worth tens of meV and has one sign; it cannot make Cr -200 against Mo +190.

The two-route reference comparison, redone in one convention — fitted references at the 0.10 A cut against the lowest corrected pure cell from the pure_elem_lowE pass:

Cr +265   Mo -129   V +91   Nb +27   Ta +28   W -25   meV/atom (fitted minus pure)

Still scattered, and pure Cr cells needed only 63 meV/atom of correction, so this is not the correction either. The mechanism, and the one I recorded in memory and on the plan page is withdrawn: a composition-linear regression over alloys does not recover elemental energies. Its intercepts are "elemental energy plus the mean mixing slope along that element's direction", and with five pure-Cr frames against thousands of Cr-containing alloys the Cr column is set by the alloys. That is a property of regression intercepts, not of MACE's volume. The right references are the corrected pure cells where they exist (six elements) with only Hf, Ti and Zr fitted — a follow-up for the hull, since a per-element constant is absorbed exactly by any expansion's point terms and does not touch ordering.

P1 died on my bug. train_ece.py wrote each m x m x m supercell's fractional coordinates into a 1 x 1 x 1 cell — 54 atoms in a 3.29 A cube — and pyeCE rejected all 4,326 structures as "lattice deformation too large". ideal_frac_of returns m; emit discarded it. Fixed; relaunched to a fresh directory.

EXPERIMENTS.md · lines 10584–10597

E151 follow-up — references from corrected pure cells, done. --pure-references holds Cr, Mo, Nb, Ta, V, W at the lowest corrected pure bcc cell from the pure_elem_lowE pass and regresses only Hf, Ti, Zr on the remainder:

Cr -9.5084   Mo -10.9154   Nb -10.2122   Ta -11.7940   V -8.9864   W -12.9257   (measured)
Hf -12.3721  Ti -7.6645    Zr -8.1505                                            (fitted)

formation energy, 4,326 frames at the 0.10 A cut:  mean +17.2  sd 83.1  min -270  max +324 meV/atom

The mean is no longer zero by construction — a set of real bcc formation energies has a mean, and this one is positive because the equation-of-state sets sample unfavourable mixtures on purpose. This file, rhea_labels_formation_strain0.10_pureref.jsonl, is the one a hull is built from. Ordering energies are unchanged by a per-element shift; the eCE training file need not be rebuilt for P1, only for any formation-energy use.

Related entries

Built with PRISMWebsite and visualizations made using Claude