Experiments · E158

Does the new energy model rank atom arrangements better than the old one?

Partly. It ranked better on unseen compositions (0.80 against 0.65) and on held-out MoNbTaVW, but the old model won on MoNbTaW (0.83 against 0.52).

In the log: P1: the embedded cluster expansion trained on the clean RHEA labels

mixedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 4 · DFT0 predictions · 2 result paragraphsEXPERIMENTS.md lines 10220–10240, lines 10242–10269, lines 10271–10294
exp E158 diagram
What E158 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E158.svg).

Results

EXPERIMENTS.md · line 10242

E158 result, in-distribution — the eCE beats the old expansion on the quantity that matters, and the comparison is clean. pyece predict on the full database, scored by the same within-composition protocol as E151's baseline, split by pyeCE's own recorded indexes:

split          n     MAE   RMSE  |  comps>=3   rho/comp   |dE| DFT   |dE| eCE     (meV/atom)
train       3,227    8.0   10.5  |    160       0.823       28.8       29.2
validation    462   10.2   13.2  |     25       0.801       24.5       25.5
test          478   10.9   14.2  |     26       0.803       33.9       36.2
baseline, all 4,326 (E151):        icet  rho 0.651, MAE-dE 14.5   MACE  rho 0.671, MAE-dE 12.0

Ranking. On compositions the model never trained on (the test split is composition-wise; E151 measured 493 of 1,834 compositions carry siblings, so a frame split would have leaked), the eCE ranks configurations at rho/comp 0.803 against icet's 0.651 and MACE's 0.671, with the ordering-energy scale right to 7%. Train is 0.823 — the gap to test is 0.02, so this is not memorisation. Its MAE on the within-composition differences is 6.3 against icet's 14.5 (pooled set). That difference — which configuration sits lower, by how much — is what Monte Carlo samples and what a transition temperature is made of.

Absolute error. RMSE 14.2 on the test split is at the labels' 12.4 noise floor, the same place icet and MACE sit. No model can be separated from perfect there, and none is.

What was withdrawn on the way here stays withdrawn: the "17x / 45x too flat" readings of E150/E151 were correction error. The old expansion is competent; the eCE is better at ranking on unseen compositions, and covers nine elements in one model without a refit.

Holdout (MoNbTaW / MoNbTaVW, 159 frames never seen at those compositions): eCE RMSE 14.9, MAE 11.4 (MoNbTaW alone MAE 6.1). The same-structure ranking against icet and MACE is running and is appended below when it lands.

EXPERIMENTS.md · line 10271

E158 result, holdout — mixed, and recorded as mixed. The 159 MoNbTaW / MoNbTaVW frames were withheld from training entirely. Ranked per system across all its frames (composition and configuration variance together, since only one composition group has three frames):

system       n  |  rho eCE   rho icet   rho MACE  |  |dE| DFT   eCE    icet   MACE
MoNbTaVW   113  |   0.570     0.238      0.415    |    16.6     7.8     7.2   255.2
MoNbTaW     46  |   0.522     0.831      0.347    |    18.1    17.6    14.2    61.1

On the quinary the eCE ranks best — 0.57 against the old expansion's 0.24 — which is the system the old expansion has always been weakest on. On the quaternary the old expansion wins outright, 0.83 to 0.52, and MoNbTaW is the alloy this project validated it on from E113 onward. The eCE's ordering scale on MoNbTaVW is half of DFT's (7.8 vs 16.6); on MoNbTaW it is right (17.6 vs 18.1). MACE's system-level spread of 255 and 61 meV/atom means its absolute energies do not share one reference across compositions; only its within-composition numbers (E151: 0.671) are meaningful, and it is not a contender here.

So the extrapolation claim is half-confirmed and half-falsified. In-distribution the eCE ranks clearly better (E158, 0.803 vs 0.651 on a composition-wise test split). Held out, it is better where the old model was bad and worse where the old model was good. Two readings, both testable: the old expansion's MACE-label training set covered MoNbTaW densely and that is what 0.83 reflects; or the eCE at 6.0 / 4.5 A cutoffs is under-resolving the quaternary, which the paper's 10 / 4 A setting (P1's deferred second run) would show. The numbers rest on 46 and 113 frames and one composition group; they are a signal, not a verdict.

The full record

This entry is written in 3 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 10220–10240

E158 — P1: the embedded cluster expansion trained on the clean RHEA labels

runs/ece_v3b_strain010: 4,326 structures at the 0.10 A strain cut, every cell size, composition-wise split (1,468 / 183 / 183 compositions), 159 MoNbTaW + MoNbTaVW frames held out in a separate test file, 3-dimensional embedding, 16-16 network, pair 6.0 A / triplet 4.5 A (the icet expansion's cutoffs, so the comparison has one variable), 400 epochs.

training loss     1.101e-4
validation loss   1.749e-4
test loss         2.016e-4      (composition-wise test split, not the holdout)

Read as mean squared error in eV²/atom², that is RMSE 10.5 / 13.2 / 14.2 meV/atom — the test figure sitting at the labels' 12.4 meV/atom noise floor, which is also where the old icet expansion sits (MAE 14.5, E151). On this set, then, the eCE is as good as the labels allow and no better than the model it was built to replace, exactly as the P2 baseline predicted the in-distribution comparison would come out.

Pending, and decisive: the held-out MoNbTaW / MoNbTaVW file — chemistry the model never saw at those compositions — and the within-composition rank correlation scored the same way as E151's baseline (icet 0.651, MACE 0.671). The loss definition is confirmed from pyeCE's own log before any of these numbers are quoted as meV.

EXPERIMENTS.md · lines 10242–10269

E158 result, in-distribution — the eCE beats the old expansion on the quantity that matters, and the comparison is clean. pyece predict on the full database, scored by the same within-composition protocol as E151's baseline, split by pyeCE's own recorded indexes:

split          n     MAE   RMSE  |  comps>=3   rho/comp   |dE| DFT   |dE| eCE     (meV/atom)
train       3,227    8.0   10.5  |    160       0.823       28.8       29.2
validation    462   10.2   13.2  |     25       0.801       24.5       25.5
test          478   10.9   14.2  |     26       0.803       33.9       36.2
baseline, all 4,326 (E151):        icet  rho 0.651, MAE-dE 14.5   MACE  rho 0.671, MAE-dE 12.0

Ranking. On compositions the model never trained on (the test split is composition-wise; E151 measured 493 of 1,834 compositions carry siblings, so a frame split would have leaked), the eCE ranks configurations at rho/comp 0.803 against icet's 0.651 and MACE's 0.671, with the ordering-energy scale right to 7%. Train is 0.823 — the gap to test is 0.02, so this is not memorisation. Its MAE on the within-composition differences is 6.3 against icet's 14.5 (pooled set). That difference — which configuration sits lower, by how much — is what Monte Carlo samples and what a transition temperature is made of.

Absolute error. RMSE 14.2 on the test split is at the labels' 12.4 noise floor, the same place icet and MACE sit. No model can be separated from perfect there, and none is.

What was withdrawn on the way here stays withdrawn: the "17x / 45x too flat" readings of E150/E151 were correction error. The old expansion is competent; the eCE is better at ranking on unseen compositions, and covers nine elements in one model without a refit.

Holdout (MoNbTaW / MoNbTaVW, 159 frames never seen at those compositions): eCE RMSE 14.9, MAE 11.4 (MoNbTaW alone MAE 6.1). The same-structure ranking against icet and MACE is running and is appended below when it lands.

EXPERIMENTS.md · lines 10271–10294

E158 result, holdout — mixed, and recorded as mixed. The 159 MoNbTaW / MoNbTaVW frames were withheld from training entirely. Ranked per system across all its frames (composition and configuration variance together, since only one composition group has three frames):

system       n  |  rho eCE   rho icet   rho MACE  |  |dE| DFT   eCE    icet   MACE
MoNbTaVW   113  |   0.570     0.238      0.415    |    16.6     7.8     7.2   255.2
MoNbTaW     46  |   0.522     0.831      0.347    |    18.1    17.6    14.2    61.1

On the quinary the eCE ranks best — 0.57 against the old expansion's 0.24 — which is the system the old expansion has always been weakest on. On the quaternary the old expansion wins outright, 0.83 to 0.52, and MoNbTaW is the alloy this project validated it on from E113 onward. The eCE's ordering scale on MoNbTaVW is half of DFT's (7.8 vs 16.6); on MoNbTaW it is right (17.6 vs 18.1). MACE's system-level spread of 255 and 61 meV/atom means its absolute energies do not share one reference across compositions; only its within-composition numbers (E151: 0.671) are meaningful, and it is not a contender here.

So the extrapolation claim is half-confirmed and half-falsified. In-distribution the eCE ranks clearly better (E158, 0.803 vs 0.651 on a composition-wise test split). Held out, it is better where the old model was bad and worse where the old model was good. Two readings, both testable: the old expansion's MACE-label training set covered MoNbTaW densely and that is what 0.83 reflects; or the eCE at 6.0 / 4.5 A cutoffs is under-resolving the quaternary, which the paper's 10 / 4 A setting (P1's deferred second run) would show. The numbers rest on 46 and 113 frames and one composition group; they are a signal, not a verdict.

Related entries

Built with PRISMWebsite and visualizations made using Claude