Can the energy model be made exact for pure elements without losing accuracy elsewhere?
Yes. Weighting the eight pure elements tenfold cut their worst error from 57.5 to 4 meV/atom; accuracy elsewhere barely moved, but ordering temperatures shifted.
In the log: Refitting the expansion with the eight elements: prediction
confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04rung 3 · kinetics0 predictions · 1 result paragraphEXPERIMENTS.md lines 4346–4436
What E80 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E80.svg).
Pre-registration
The pre-registration, as written
E80 — Refitting the expansion with the eight elements: prediction
The repair E72 named and deferred. The expansion is fitted to mixing energies against
shared-lattice elemental references, so a pure element's target is exactly zero - it has
nothing to mix. Evaluated on its own convention the current fit says otherwise:
Hf +4.5 Mo -57.5 Nb +3.5 Ta +14.8 Ti +13.4 V -17.6 W -49.4 Zr +4.7
which reproduces the E72 errors to the decimal and confirms the diagnosis from the other
direction. Adding the eight corners costs no new data: the cluster vectors come from
icet and the targets are zero by definition. The stored design matrix makes the refit a
lasso on 1961 rows by 344 columns.
Predicted:
The corners are pinned - all eight within a few meV/atom of zero, where Mo and W are
now out by 57 and 49.
The interior barely moves. Against MACE on the 120-composition relaxation survey the
present fit has a mean absolute error of 19.6 meV/atom and a bias of +3.9; both stay
within about 3 meV/atom. Eight rows against 1953 is a small perturbation, but they sit at
extreme leverage, so this is the prediction that can fail.
CV RMSE stays near 5.67 meV/atom, rising by less than 1.
The screen's pure-element false positives disappear without the leverage term.E72
patched around this failure with a composition-dependent sigma; if the refit is the real
repair, the raw driving force at a pure element should come back near zero on its own.
Falsified if the interior error degrades by more than 5 meV/atom, which would mean the
corners can only be pinned by spending accuracy where the search actually lives - a trade
worth knowing about and not worth making silently.
Not run on a GPU, and the reason is worth stating. The request was to refit on the A100.
There is nothing here for one: the expansion is a linear model whose design matrix is already
on disk, so the fit is milliseconds of CPU. What a GPU accelerates in this project is MACE
inference - the 136-phase hull, the relaxations, the barriers - and the refit consumes none
of it. Renting an A100 for this would be renting it to sit idle.
Outcome. The corners are pinned, the interior is untouched, and every ordering
temperature in this file is now stale.
The trade was measured rather than guessed, by weighting the eight corners against the 1953
ordinary structures and watching both ends move:
corner weight
CV RMSE
worst corner
interior MAE
bias
0 (the fit in use)
5.67
57.5
18.7
+1.2
1
5.92
25.7
19.0
+1.4
3
5.79
12.2
19.2
+1.6
10 (deployed)
5.68
4.3
19.1
+1.8
30
5.46
1.5
19.2
+2.0
100
4.90
0.5
19.3
+2.3
Prediction 1 failed at weight 1 and holds at weight 10. Eight rows against 1953 is too
little leverage to pin anything - molybdenum only came back from -58 to -26. The prediction
named an outcome without naming what it would cost to reach it, which is half a prediction.
At weight 10 the corners are Hf 0, Mo -4, Nb 0, Ta +1, Ti +1, V -1, W -3, Zr 0: worst 4
meV/atom, inside the expansion's own cross-validation error, where it was ten times it.
Predictions 2 and 3 hold. Interior mean absolute error against MACE on the 120
relaxation-survey compositions moves 19.6 to 19.9, spread 24.4 to 25.1, bias +3.9 to +4.3.
CV RMSE 5.67 to 5.68. The corners were pinned for four tenths of a meV/atom.
The CV column past weight 30 is not a fair comparison and should not be read as one: the
training set is padded with duplicate corner rows that any fit reproduces, so the number
improves by being asked an easier question.
Prediction 4 holds, and the leverage machinery retired itself. The raw driving force at
a pure element is now -4 to +1 meV/atom where it was -58, with no uncertainty term needed.
And sigma at the corners fell from about 130 to 38 unprompted, because leverage there
dropped from 3.3-5.9 to 0.30: the corners stopped being extrapolation the moment they became
training points. Interior leverage is unchanged - median 0.173 to 0.169 - so the sigma
calibration of E72 carries over without refitting.
What this costs, and it is the thing to read in the morning. Driving forces move by 4
meV/atom or less and no ranking changes. Ordering temperatures move by up to 265 K:
composition
T_od before
T_od after
scatter before
after
MoNbTaW
501
635
+/-182
+/-79
Mo0.62 Ta0.38
1077
1342
+/-202
+/-45
Ta0.39 Mo0.34 W0.18 Nb0.08
989
773
+/-111
+/-76
Mo0.33 Nb0.15 Ta0.21 W0.31
713
805
+/-27
+/-62
Every ordering temperature recorded in this file was measured with the prior fit and is
withdrawn as a number, though not as a conclusion: the driving forces that decided the
rankings are unchanged. The scatter halving on three of four is the argument that the new
values are better and not merely different - the prior fit's corner pathology was feeding
noise into a Monte Carlo that never goes near a corner. E74's verdicts rest on T_od and must
be re-measured before they are quoted again.
The prior fit is kept at data/ce_8element_prior.*. A regression test now asserts every
corner is inside the fit's own CV error. 222 passing.
Results
EXPERIMENTS.md · line 4382
Outcome. The corners are pinned, the interior is untouched, and every ordering
temperature in this file is now stale.
The full record
EXPERIMENTS.md · lines 4346–4436
E80 — Refitting the expansion with the eight elements: prediction
The repair E72 named and deferred. The expansion is fitted to mixing energies against
shared-lattice elemental references, so a pure element's target is exactly zero - it has
nothing to mix. Evaluated on its own convention the current fit says otherwise:
Hf +4.5 Mo -57.5 Nb +3.5 Ta +14.8 Ti +13.4 V -17.6 W -49.4 Zr +4.7
which reproduces the E72 errors to the decimal and confirms the diagnosis from the other
direction. Adding the eight corners costs no new data: the cluster vectors come from
icet and the targets are zero by definition. The stored design matrix makes the refit a
lasso on 1961 rows by 344 columns.
Predicted:
The corners are pinned - all eight within a few meV/atom of zero, where Mo and W are
now out by 57 and 49.
The interior barely moves. Against MACE on the 120-composition relaxation survey the
present fit has a mean absolute error of 19.6 meV/atom and a bias of +3.9; both stay
within about 3 meV/atom. Eight rows against 1953 is a small perturbation, but they sit at
extreme leverage, so this is the prediction that can fail.
CV RMSE stays near 5.67 meV/atom, rising by less than 1.
The screen's pure-element false positives disappear without the leverage term.E72
patched around this failure with a composition-dependent sigma; if the refit is the real
repair, the raw driving force at a pure element should come back near zero on its own.
Falsified if the interior error degrades by more than 5 meV/atom, which would mean the
corners can only be pinned by spending accuracy where the search actually lives - a trade
worth knowing about and not worth making silently.
Not run on a GPU, and the reason is worth stating. The request was to refit on the A100.
There is nothing here for one: the expansion is a linear model whose design matrix is already
on disk, so the fit is milliseconds of CPU. What a GPU accelerates in this project is MACE
inference - the 136-phase hull, the relaxations, the barriers - and the refit consumes none
of it. Renting an A100 for this would be renting it to sit idle.
Outcome. The corners are pinned, the interior is untouched, and every ordering
temperature in this file is now stale.
The trade was measured rather than guessed, by weighting the eight corners against the 1953
ordinary structures and watching both ends move:
corner weight
CV RMSE
worst corner
interior MAE
bias
0 (the fit in use)
5.67
57.5
18.7
+1.2
1
5.92
25.7
19.0
+1.4
3
5.79
12.2
19.2
+1.6
10 (deployed)
5.68
4.3
19.1
+1.8
30
5.46
1.5
19.2
+2.0
100
4.90
0.5
19.3
+2.3
Prediction 1 failed at weight 1 and holds at weight 10. Eight rows against 1953 is too
little leverage to pin anything - molybdenum only came back from -58 to -26. The prediction
named an outcome without naming what it would cost to reach it, which is half a prediction.
At weight 10 the corners are Hf 0, Mo -4, Nb 0, Ta +1, Ti +1, V -1, W -3, Zr 0: worst 4
meV/atom, inside the expansion's own cross-validation error, where it was ten times it.
Predictions 2 and 3 hold. Interior mean absolute error against MACE on the 120
relaxation-survey compositions moves 19.6 to 19.9, spread 24.4 to 25.1, bias +3.9 to +4.3.
CV RMSE 5.67 to 5.68. The corners were pinned for four tenths of a meV/atom.
The CV column past weight 30 is not a fair comparison and should not be read as one: the
training set is padded with duplicate corner rows that any fit reproduces, so the number
improves by being asked an easier question.
Prediction 4 holds, and the leverage machinery retired itself. The raw driving force at
a pure element is now -4 to +1 meV/atom where it was -58, with no uncertainty term needed.
And sigma at the corners fell from about 130 to 38 unprompted, because leverage there
dropped from 3.3-5.9 to 0.30: the corners stopped being extrapolation the moment they became
training points. Interior leverage is unchanged - median 0.173 to 0.169 - so the sigma
calibration of E72 carries over without refitting.
What this costs, and it is the thing to read in the morning. Driving forces move by 4
meV/atom or less and no ranking changes. Ordering temperatures move by up to 265 K:
composition
T_od before
T_od after
scatter before
after
MoNbTaW
501
635
+/-182
+/-79
Mo0.62 Ta0.38
1077
1342
+/-202
+/-45
Ta0.39 Mo0.34 W0.18 Nb0.08
989
773
+/-111
+/-76
Mo0.33 Nb0.15 Ta0.21 W0.31
713
805
+/-27
+/-62
Every ordering temperature recorded in this file was measured with the prior fit and is
withdrawn as a number, though not as a conclusion: the driving forces that decided the
rankings are unchanged. The scatter halving on three of four is the argument that the new
values are better and not merely different - the prior fit's corner pathology was feeding
noise into a Monte Carlo that never goes near a corner. E74's verdicts rest on T_od and must
be re-measured before they are quoted again.
The prior fit is kept at data/ce_8element_prior.*. A regression test now asserts every
corner is inside the fit's own CV error. 222 passing.
Related entries
E72 — A pure element screened as stable, and the relaxation model was not to blame
E74 — The generator's own candidates up the ladder: better than the reference on the hull,…