Do the energy model's training labels get the molybdenum–tantalum mixing energy right?
Partly. DFT gave −89.8 meV/atom: near the labels referenced to pure elements (−73.7), far from the fitted-reference labels the model used (−22.9).
In the log: Mo₂₇Ta₂₇: the basin against the literature, and a test of the labels themselves
recordedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 11481–11534, lines 11548–11558
What E180 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E180.svg).
Results
EXPERIMENTS.md · line 11548
E180 result — between the two branches, and on the pure-referenced side. JOB DONE in
21 iterations; E/atom −7422.32794 eV at Vegard; references Mo, Ta (refs_v2/):
E_form(DFT, QE) = −89.8 meV/atom. Branch A's window [−150, −100] missed by 10; branch
B's [−40, −5] missed by 50. Against the three stated expectations: literature random (LDA
MBCE) −127: 37 off; RHEA's pure-referenced label −73.7: 16 off; the regression label
−22.9: 67 off. Inside the measured mixing enthalpy −114 ± 26. So: the regression-reference
label is falsified as a label; the pure-referenced convention is the right one; and the
16 meV between the QE ideal-site cell and RHEA's corrected cell is the size of the rattle
correction itself (+86 meV on a 0.32 Å-rattled cell) — the correction's error, not a code
offset (Mo and Ta are V-free, no cutoff caveat). The icet CE (−11.5) and eCE v3b (−23.6)
were 4–8× off here, as E177 already showed for the basin.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 11481–11534
E180 — Mo₂₇Ta₂₇: the basin against the literature, and a test of the labels themselves
The external numbers. Blum & Zunger (PRB 69, 214202 / PRB 72, 020104; LDA mixed-basis
cluster expansion): the fully random Mo–Ta at x = 0.5 has ΔH_mix = −127 meV/atom; the
measured mixing enthalpy is −114 ± 26. Widom (arXiv:1306.5043, PBE): ordered MoTa
(oA12/B2) −186 meV/atom, Mo–Ta the strongest-bonding pair of the MoNbTaW binaries.
The internal numbers. icet rung 0 for Mo–Ta 50/50: e_bcc −11.5 meV/atom, T_order
1782 K (above the window: a pass), p 0.997. eCE for the seed-0 Mo₂₇Ta₂₇ decoration:
−23.6. And the finding that changes what this cell tests: RHEA's own 54-atom random
Mo–Ta 50/50 cell carries a label of −22.9 meV/atom in my corrected label set (the eCE
gives it −9.9; the three 16-atom Mo–Ta cells carry −133, −29, −21). The model is faithful
to its labels. So either (a) the label pipeline — rattle correction by MACE difference,
volume relaxation, strain cut, pure-cell referencing — has removed ~100 meV/atom of mixing
energy from random Mo–Ta cells, or (b) RHEA's PBE-VASP random Mo–Ta really is at −23 and
the LDA-CE literature is 5× away, which no code difference explains.
Cell. Mo₂₇Ta₂₇, seed-0 decoration on ideal sites, Vegard a = 3.2460 Å (L 9.7380),
E176's settings, single SCF; references Mo and Ta from refs_v2/.
Predictions — two branches, one of which this cell falsifies.
Branch A (literature): E_form(DFT, QE) in [−150, −100]. Then RHEA's raw VASP energy for
its random cell must also be near −120 and my correction took it to −23: the label
pipeline is the defect, every "compression" measured this week is largely a label
artefact, and E172/E176/E177 must be re-read against raw labels.
Branch B (labels): E_form(DFT, QE) in [−40, −5]. Then the labels are right, the QE cell
agrees with RHEA's VASP, and the literature random value is not comparable (LDA, MBCE
fit) — the cheap rungs are consistent with DFT at this composition and E177's −71 was
the Ti/Hf/V admixture, not the Mo–Ta basin.
Anything between −100 and −40 is neither and is stated as such.
The raw-versus-corrected label for RHEA's Mo–Ta cell is read before DFT finishes and
recorded beside this, so the DFT verdict lands on a stated expectation.
Read before DFT — and it is my pipeline. RHEA index 6369 (54-atom random Mo–Ta 50/50,
bcc_alloys): raw E_DFT −11.3426 eV/atom, rattle+volume correction +0.0858 eV,
E_lattice −11.4284. Referenced to corrected pure cells: e_formation = −73.7 meV/atom
(rhea_labels_formation_strain0.10_pureref.jsonl). Referenced to the regression
intercepts: −22.9 (rhea_labels_formation_strain0.10.jsonl). The deployed model,
ece_v3b_strain010, was fitted on the second file — the training property for this cell
is −22.9, and the eCE reproduces it. So:
The regression references absorb most of the mean mixing energy into the "elements",
and formation energies referenced to them are compressed toward zero by construction.
That is the 0.59 holdout slope, and it is the 1.9× / 2.2× / 4.0× of E172, E176 and E177 —
not the model extrapolating, as E176 and E177 concluded. Those readings are
withdrawn in part: the model may still compress on unseen compositions, but that
cannot be measured until it is trained on labels whose zero is the elements.
The memory note "regression references ≠ elemental energies → use corrected pure cells"
existed; --pure-references was built; the model that went into service was the one
fitted before it. A rule for the record: the label file's basename is part of the
model's name from now on.
E180's DFT now lands on three stated expectations: literature −127 (LDA, random), the
pure-referenced label −74 (PBE-VASP, RHEA's own random cell), the regression label −23.
Same-code agreement with RHEA is the −74; the QE cell should sit within ±12 of it if the
two PBE-PAW codes agree on Mo–Ta as ACWF says they should.
EXPERIMENTS.md · lines 11548–11558
E180 result — between the two branches, and on the pure-referenced side. JOB DONE in
21 iterations; E/atom −7422.32794 eV at Vegard; references Mo, Ta (refs_v2/):
E_form(DFT, QE) = −89.8 meV/atom. Branch A's window [−150, −100] missed by 10; branch
B's [−40, −5] missed by 50. Against the three stated expectations: literature random (LDA
MBCE) −127: 37 off; RHEA's pure-referenced label −73.7: 16 off; the regression label
−22.9: 67 off. Inside the measured mixing enthalpy −114 ± 26. So: the regression-reference
label is falsified as a label; the pure-referenced convention is the right one; and the
16 meV between the QE ideal-site cell and RHEA's corrected cell is the size of the rattle
correction itself (+86 meV on a 0.32 Å-rattled cell) — the correction's error, not a code
offset (Mo and Ta are V-free, no cutoff caveat). The icet CE (−11.5) and eCE v3b (−23.6)
were 4–8× off here, as E177 already showed for the basin.
Related entries
E176 — the second rung-4 verdict: a refractory-centred cell
E172 — the first rung-4 verdict: DFT vs the cheap rung on one 54-atom cell