At a composition the search itself found, does DFT agree with the cheap models?
No. DFT gave −71.1 meV/atom against −17.9 and −10.9 from the two models; most of the gap was later traced to the labels' references.
In the log: the third rung-4 cell, at a search find
falsifiedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 11368–11391, lines 11428–11467
What E177 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E177.svg).
Pre-registration
The pre-registration, as written
E177 — the third rung-4 cell, at a search find
E175's refit is deferred: two QE points cannot separate an offset from a scale, and the
slope is already measured. What the flywheel needs is cells at the search's own finds,
where the compression is paid out. The search's rung 0 is the icet CE on ce_12element
(not the eCE), so this cell tests both. The finds are not recorded by composition in the
run logs, so one was taken from the same screen: 300 random compositions over the eight
pseudopotential elements, best qualifying hit (p ≥ 0.832) at L1 ≥ 0.15 from every RHEA
training composition and from MoNbTaW: Mo 0.488 Ta 0.328 Ti 0.098 Hf 0.044 V 0.042 —
the Mo–Ta basin the fly and the CEM keep returning to. icet rung 0: e_bcc −10.9 meV/atom,
drive_cold −67.2, drive_hot −162, σ 36.6, T_order 709 K, p 0.910. Cell: Mo₂₇Ta₁₈Ti₅Hf₂V₂
(multinomial rounding), seed-0 decoration, Vegard a = 3.2419 Å, E176's settings, single
SCF. eCE E_lattice = −17.9 meV/atom; references from refs_v2/.
Predictions.
The compression holds at the find: E_form(DFT) = 2.0 × eCE ± 10 → in [−46, −26]
(E172 and E176 gave 1.90 and 2.24). Falsified if DFT is within ±8 of −17.9 — then the
compression is not uniform and E176's centre-of-data case was a coincidence; or if DFT
is below −50 — then the cheap rungs under-read this basin by more than the scale.
The icet CE (the rung that paid for this find) is off by more than 20 meV/atom
(−10.9 vs DFT): it is the model the memory records as wrong at the corners and it has
the smaller magnitude here. If it lands within 10, the search's own rung is better at
its own finds than the eCE — which would reverse the plan to make the eCE the bottom rung.
Converges in < 40 iterations (5 species).
Results
EXPERIMENTS.md · line 11428
E177 result — prediction 1 falsified downward: the cheap rungs under-read the search's
own basin by 4×. JOB DONE in 36 iterations (prediction 3 confirmed; the SCF tail crept
from 1e-6 to 1e-7 over ten iterations — Mo–Ta's dense Fermi surface, not sloshing).
E/atom −6463.52661 eV at Vegard; references Hf/Mo/Ta/Ti/V from refs_v2/, all bracketed:
Prediction 1: falsified — below −50, by the branch written in advance: "the cheap
rungs under-read this basin by more than the scale".
Prediction 2: confirmed — the icet CE, the rung that paid for this find, is off by
60 meV/atom, three times the bar.
Three cells now:
cell eCE icet DFT DFT/eCE where
E172 Hf/Zr-rich +51.2 — +97.2 1.9 thin edge
E176 MoNbTaW −21.9 — −49.1 2.2 data centre
E177 Mo–Ta basin find −17.9 −10.9 −71.1 4.0 the search's basin
Not an offset, not a constant scale: a composition-dependent under-reading, worst where
the search goes. The sign is the consoling part — the finds are more favourable than
they were paid for, not less. The uncomfortable part is the ordering temperature: rung 1's
T_order for this composition (709 K, inside the window, so a fail) is proportional to the
ordering energy the same models supply; under-read by 4×, the true ordering temperature
is well above the window, which is a pass by the ladder's own rule ("ordered throughout,
a different material"). The cheap rungs are not just miscalibrating the reward; they are
placing the transition on the wrong side of 1000 K for the search's best basin.
Branch. (a) The flywheel refit must be local, not a scale — the eCE's species embedding
can carry per-composition corrections, but only if the QE cells sit in its training set with
the per-source term, and three points are still too few. (b) Before that, the most
load-bearing claim tonight rests on one decoration: E179 replicates E177 with the
seed-1 decoration of Mo₂₇Ta₁₈Ti₅Hf₂V₂ at the same Vegard a (19 of 54 sites coincide with
seed 0; eCE scores it −11.9 vs seed 0's −17.9 — a 6 meV decoration spread on the model's
side, recorded before DFT). Prediction: E_form within the decoration noise floor,
|ΔE| ≤ 12.4 meV/atom of −71.1; if the two decorations
differ by more than 25, ordering within the cell (a lucky Mo–Ta arrangement) is doing the
work and the composition-level number needs several decorations. (c) The icet CE's 60 meV
miss at its own find is the number that should replace "8.1 meV vs DFT" in every summary
of the bottom rung: that figure was a fit residual on RHEA's distribution; this is the error
where the search lives.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 11368–11391
E177 — the third rung-4 cell, at a search find
E175's refit is deferred: two QE points cannot separate an offset from a scale, and the
slope is already measured. What the flywheel needs is cells at the search's own finds,
where the compression is paid out. The search's rung 0 is the icet CE on ce_12element
(not the eCE), so this cell tests both. The finds are not recorded by composition in the
run logs, so one was taken from the same screen: 300 random compositions over the eight
pseudopotential elements, best qualifying hit (p ≥ 0.832) at L1 ≥ 0.15 from every RHEA
training composition and from MoNbTaW: Mo 0.488 Ta 0.328 Ti 0.098 Hf 0.044 V 0.042 —
the Mo–Ta basin the fly and the CEM keep returning to. icet rung 0: e_bcc −10.9 meV/atom,
drive_cold −67.2, drive_hot −162, σ 36.6, T_order 709 K, p 0.910. Cell: Mo₂₇Ta₁₈Ti₅Hf₂V₂
(multinomial rounding), seed-0 decoration, Vegard a = 3.2419 Å, E176's settings, single
SCF. eCE E_lattice = −17.9 meV/atom; references from refs_v2/.
Predictions.
The compression holds at the find: E_form(DFT) = 2.0 × eCE ± 10 → in [−46, −26]
(E172 and E176 gave 1.90 and 2.24). Falsified if DFT is within ±8 of −17.9 — then the
compression is not uniform and E176's centre-of-data case was a coincidence; or if DFT
is below −50 — then the cheap rungs under-read this basin by more than the scale.
The icet CE (the rung that paid for this find) is off by more than 20 meV/atom
(−10.9 vs DFT): it is the model the memory records as wrong at the corners and it has
the smaller magnitude here. If it lands within 10, the search's own rung is better at
its own finds than the eCE — which would reverse the plan to make the eCE the bottom rung.
Converges in < 40 iterations (5 species).
EXPERIMENTS.md · lines 11428–11467
E177 result — prediction 1 falsified downward: the cheap rungs under-read the search's
own basin by 4×. JOB DONE in 36 iterations (prediction 3 confirmed; the SCF tail crept
from 1e-6 to 1e-7 over ten iterations — Mo–Ta's dense Fermi surface, not sloshing).
E/atom −6463.52661 eV at Vegard; references Hf/Mo/Ta/Ti/V from refs_v2/, all bracketed:
Prediction 1: falsified — below −50, by the branch written in advance: "the cheap
rungs under-read this basin by more than the scale".
Prediction 2: confirmed — the icet CE, the rung that paid for this find, is off by
60 meV/atom, three times the bar.
Three cells now:
cell eCE icet DFT DFT/eCE where
E172 Hf/Zr-rich +51.2 — +97.2 1.9 thin edge
E176 MoNbTaW −21.9 — −49.1 2.2 data centre
E177 Mo–Ta basin find −17.9 −10.9 −71.1 4.0 the search's basin
Not an offset, not a constant scale: a composition-dependent under-reading, worst where
the search goes. The sign is the consoling part — the finds are more favourable than
they were paid for, not less. The uncomfortable part is the ordering temperature: rung 1's
T_order for this composition (709 K, inside the window, so a fail) is proportional to the
ordering energy the same models supply; under-read by 4×, the true ordering temperature
is well above the window, which is a pass by the ladder's own rule ("ordered throughout,
a different material"). The cheap rungs are not just miscalibrating the reward; they are
placing the transition on the wrong side of 1000 K for the search's best basin.
Branch. (a) The flywheel refit must be local, not a scale — the eCE's species embedding
can carry per-composition corrections, but only if the QE cells sit in its training set with
the per-source term, and three points are still too few. (b) Before that, the most
load-bearing claim tonight rests on one decoration: E179 replicates E177 with the
seed-1 decoration of Mo₂₇Ta₁₈Ti₅Hf₂V₂ at the same Vegard a (19 of 54 sites coincide with
seed 0; eCE scores it −11.9 vs seed 0's −17.9 — a 6 meV decoration spread on the model's
side, recorded before DFT). Prediction: E_form within the decoration noise floor,
|ΔE| ≤ 12.4 meV/atom of −71.1; if the two decorations
differ by more than 25, ordering within the cell (a lucky Mo–Ta arrangement) is doing the
work and the composition-level number needs several decorations. (c) The icet CE's 60 meV
miss at its own find is the number that should replace "8.1 meV vs DFT" in every summary
of the bottom rung: that figure was a fit residual on RHEA's distribution; this is the error
where the search lives.
Related entries
E175 — named in the log, no entry of its own
E176 — the second rung-4 verdict: a refractory-centred cell
E172 — the first rung-4 verdict: DFT vs the cheap rung on one 54-atom cell
E179 — result — prediction confirmed: the −71 is the composition's, not the decoration's