Does the new energy model reproduce published ordering temperatures for nine alloys?
Withdrawn. It scored 0 of 10 within 200 K, but the test was too coarse and its temperatures were half the true ones.
In the log: rung 1 from v5, scored against the published ordering table
falsifiedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 11992–12008, lines 12010–12019, lines 12028–12054
What E190 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E190.svg).
Pre-registration
The pre-registration, as written
E190 — rung 1 from v5, scored against the published ordering table
forager/ladder/ece_odt.py::transition_ece asks v5 for an ordering temperature by pyeCE
canonical Monte Carlo (annealed 1800 → 300 K, 12 states, 4×4×4 = 128 sites, T_c from the
heat-capacity peak with a parabolic refinement; sigma = half-width at 70%, floored at the
sweep step) and returns what rungs.transition returns. Validation before it replaces
anything: every distinct composition in published_odt.PUBLISHED_ODT (Sobieraj 2020,
Fernández-Caballero 2017, Woodgate 2023, Kim & Widom 2023 — all DFT-CE + MC, the same
quantity by the same route), one checkpoint each, load-gated.
Predictions. (Two entries — WAK23 MoNbTa, KW23 MoNbTaW — report no transition; they are
scored separately: v5 should find none above 300 K or a weak one, and either is stated.)
≥ 7 of the 10 numeric entries within ±200 K of the published value (the table's own
method spread is 100 K; a 128-site cell and 12 states add ~±75 K of resolution). 2. Falsified
at ≤ 4 of 10: v5's ordering energies do not transfer across the table and rung 1 stays on
hold — then the failure pattern (which chemistries) is the finding. 3. The Mo–Ta-containing
entries (KW23, FC17) land below the icet CE's numbers by ≥ 2× — E188's result generalises.
Wall time ≤ 10 min per composition at 4 threads.
E190 interim (15:29) — seven of seven low, chemistry-independent. At the validation
resolution (128 sites, 12 states, 60 samples): CrTaTiVW ≤ 300 (pub. 900/1000), TaTiVW ≤ 300
(500), CrTaTiW ≤ 300 (500), CrTaVW 435 (1200/1300), CrTiVW 425 (700), CrTaTiV ≤ 300 (600),
MoNbTaVW ≤ 300 (750, 742). The last is the decisive one: a refractory-centred alloy whose
formation energies v5 gets within 10 meV/atom, and it still reads "Cv still rising at
300 K". Two readings, one test between them: (i) v5's ordering energies — the
differences between decorations at fixed composition, a much smaller quantity than the
formation energy — are weak across the board; (ii) the validation sweep is too coarse to
grow a heat-capacity peak (E188's clean Mo–Ta peak took 432 sites and 100 samples on the
strongest-ordering pair in the table).
Results
EXPERIMENTS.md · line 12028
E190 result (15:31) — prediction 1 FALSIFIED, prediction 2 fires. All nine compositions
at the validation resolution (runs/e190_odt_validation/table.json):
alloy published (K) v5 (K) read
CrTaTiVW 900 / 1000 ≤ 300 still rising at the coldest state
TaTiVW 500 ≤ 300 censored
CrTaTiW 500 ≤ 300 censored
CrTaVW 1200 / 1300 435 peak; 800 low
CrTiVW 700 425 peak; 275 low
CrTaTiV 600 ≤ 300 censored
MoNbTaVW 750 / 742 ≤ 300 censored — the DFT-validated chemistry
MoNbTa none reported 583 a peak where WAK23 reports none
MoNbTaW none reported ≤ 300 consistent with "none"
The script's count "2 of 10 within 200 K" is the two censored ≤ 300 reads against published
500 K — not hits. Scored honestly: 0 of 10. Every v5 temperature is low, by 275 K to
≥ 900 K, in every chemistry including the one whose formation energies v5 reproduces to
10 meV/atom. So rung 1 does not transfer from v5 as fitted and as swept here. Which of
the two it is — v5's within-composition ordering energies being weak (RHEA's ordered cells
are 2–16 atoms; the fit's ρ 0.80 on them is a ranking, not a scale), or the 128-site /
60-sample sweep failing to grow a peak that a 432-site / 100-sample sweep grows (E188) — is
what E190b decides, and E188's Mo–Ta 500 K stands only if E190b shows MoNbTaVW's peak at
the finer resolution. Prediction 3 (Mo–Ta entries ≥ 2× below icet) is moot: v5 is below
everything. Prediction 4 (≤ 10 min/composition) held (~7 min each). Until E190b: rung 1
stays on the icet CE with its known 3.5× excess on Mo–Ta, and no ladder verdict is
re-scored on v5's T_c — including E188's flip of the Mo–Ta basin, which is now
"provisional, pending E190b" rather than a finding.
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 11992–12008
E190 — rung 1 from v5, scored against the published ordering table
forager/ladder/ece_odt.py::transition_ece asks v5 for an ordering temperature by pyeCE
canonical Monte Carlo (annealed 1800 → 300 K, 12 states, 4×4×4 = 128 sites, T_c from the
heat-capacity peak with a parabolic refinement; sigma = half-width at 70%, floored at the
sweep step) and returns what rungs.transition returns. Validation before it replaces
anything: every distinct composition in published_odt.PUBLISHED_ODT (Sobieraj 2020,
Fernández-Caballero 2017, Woodgate 2023, Kim & Widom 2023 — all DFT-CE + MC, the same
quantity by the same route), one checkpoint each, load-gated.
Predictions. (Two entries — WAK23 MoNbTa, KW23 MoNbTaW — report no transition; they are
scored separately: v5 should find none above 300 K or a weak one, and either is stated.)
≥ 7 of the 10 numeric entries within ±200 K of the published value (the table's own
method spread is 100 K; a 128-site cell and 12 states add ~±75 K of resolution). 2. Falsified
at ≤ 4 of 10: v5's ordering energies do not transfer across the table and rung 1 stays on
hold — then the failure pattern (which chemistries) is the finding. 3. The Mo–Ta-containing
entries (KW23, FC17) land below the icet CE's numbers by ≥ 2× — E188's result generalises.
Wall time ≤ 10 min per composition at 4 threads.
EXPERIMENTS.md · lines 12010–12019
E190 interim (15:29) — seven of seven low, chemistry-independent. At the validation
resolution (128 sites, 12 states, 60 samples): CrTaTiVW ≤ 300 (pub. 900/1000), TaTiVW ≤ 300
(500), CrTaTiW ≤ 300 (500), CrTaVW 435 (1200/1300), CrTiVW 425 (700), CrTaTiV ≤ 300 (600),
MoNbTaVW ≤ 300 (750, 742). The last is the decisive one: a refractory-centred alloy whose
formation energies v5 gets within 10 meV/atom, and it still reads "Cv still rising at
300 K". Two readings, one test between them: (i) v5's ordering energies — the
differences between decorations at fixed composition, a much smaller quantity than the
formation energy — are weak across the board; (ii) the validation sweep is too coarse to
grow a heat-capacity peak (E188's clean Mo–Ta peak took 432 sites and 100 samples on the
strongest-ordering pair in the table).
EXPERIMENTS.md · lines 12028–12054
E190 result (15:31) — prediction 1 FALSIFIED, prediction 2 fires. All nine compositions
at the validation resolution (runs/e190_odt_validation/table.json):
alloy published (K) v5 (K) read
CrTaTiVW 900 / 1000 ≤ 300 still rising at the coldest state
TaTiVW 500 ≤ 300 censored
CrTaTiW 500 ≤ 300 censored
CrTaVW 1200 / 1300 435 peak; 800 low
CrTiVW 700 425 peak; 275 low
CrTaTiV 600 ≤ 300 censored
MoNbTaVW 750 / 742 ≤ 300 censored — the DFT-validated chemistry
MoNbTa none reported 583 a peak where WAK23 reports none
MoNbTaW none reported ≤ 300 consistent with "none"
The script's count "2 of 10 within 200 K" is the two censored ≤ 300 reads against published
500 K — not hits. Scored honestly: 0 of 10. Every v5 temperature is low, by 275 K to
≥ 900 K, in every chemistry including the one whose formation energies v5 reproduces to
10 meV/atom. So rung 1 does not transfer from v5 as fitted and as swept here. Which of
the two it is — v5's within-composition ordering energies being weak (RHEA's ordered cells
are 2–16 atoms; the fit's ρ 0.80 on them is a ranking, not a scale), or the 128-site /
60-sample sweep failing to grow a peak that a 432-site / 100-sample sweep grows (E188) — is
what E190b decides, and E188's Mo–Ta 500 K stands only if E190b shows MoNbTaVW's peak at
the finer resolution. Prediction 3 (Mo–Ta entries ≥ 2× below icet) is moot: v5 is below
everything. Prediction 4 (≤ 10 min/composition) held (~7 min each). Until E190b: rung 1
stays on the icet CE with its known 3.5× excess on Mo–Ta, and no ladder verdict is
re-scored on v5's T_c — including E188's flip of the Mo–Ta basin, which is now
"provisional, pending E190b" rather than a finding.
Related entries
E188 — the ordering temperature from v5, against the published Mo–Ta number