Was the low ordering temperature a coarse-simulation artefact or a real model weakness?
Withdrawn. The finer run found a peak at 367 K; doubled for the sampler's error it is 734 K, on the published 745 K.
In the log: the discriminator
withdrawnDate 2026-09-21 20:2x, as written in the logrung 4 · DFT0 predictions · 3 result paragraphsEXPERIMENTS.md lines 12020–12026, lines 12074–12095, lines 14968–14977
What E190b did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E190b.svg).
Results
EXPERIMENTS.md · line 12020
E190b — the discriminator. MoNbTaVW at E188's resolution: 6×6×6 (432 sites), 24 states
1800 → 300 K, 100 uncorrelated samples per state (max 1000), annealed. Prediction: if the
sweep was the problem, a peak appears in [550, 950] K; if it still reads ≤ 300 K, reading
(i) holds and rung 1 cannot come from v5 as fitted — the within-composition ranking (ρ 0.80
on RHEA's ordered cells) would have to be strengthened with ordered DFT cells before any
T_c is trusted, and E188's 500 K for Mo–Ta stands only because Mo–Ta's ordering energy is
large enough to survive a weak model. Queued behind E190 (same cores).
EXPERIMENTS.md · line 12074
E190b result (15:48) — both effects, separated. MoNbTaVW at E188's resolution (432
sites, 24 states 1800 → 300 K, 100 samples/state): T_c = 367 ± 65 K, a clear heat-capacity
peak (Cv 3.64 k_B), NN SRO −0.366 at 300 K — where the 128-site / 60-sample sweep had read
"still rising at 300 K". Published: 750 (FC17), 742 (WAK23). Neither pre-written branch
holds as written: the peak appears (the coarse sweep censored it: resolution effect, real)
but at half the published temperature (model effect, real). Reading, with E188 beside
it: v5's ordering temperatures come out at ~0.5× the published (Mo–Ta 500 vs 600–1000;
MoNbTaVW 367 vs 745; CrTaVW 435 vs 1250 even coarsely) — the within-composition energy
scale is weak by about 2×, while the formation energy scale is right to 10 meV/atom.
The two are different quantities: formation energy is set by the composition-wise mean the
54-atom random cells pin down; the ordering energy is the spread between decorations, which
RHEA's 2–16-atom ordered cells carry with their own rattle-correction noise (12 meV floor
against a 29 meV spread). The icet CE's excess (Mo–Ta 1782) and v5's deficit bracket the
truth from opposite sides.
Consequences: (1) rung 1 stays on hold; (2) any v5 sweep is run at ≥ 432 sites and
≥ 100 samples — the 128-site validation numbers are withdrawn as measurements (kept as
the record of a resolution failure); (3) E188's Mo–Ta 500 K is provisional in value but
robust in sign: a 2× correction puts it at ~1000 K, the published band is 600–1000 — every
estimate except the icet CE's 1782 keeps Mo–Ta's transition inside the window, so the
basin's flip from pass to fail stands; (4) branch → task 23: measure the ordering-energy
scale directly with DFT — random vs B2-type decorations of one composition in 16-atom cells,
minutes each — and feed those cells to v6.
EXPERIMENTS.md · line 14968
E190b on v5 / v6a / v6b, re-read on the physical axis (2026-09-21 20:2x). The sweeps
were 1800 → 300 K nominal = 3600 → 600 K physical. v5 MoNbTaVW: variance peak 367 K nominal
= 734 K physical, inside the sweep (dE/dT censored at the floor); α₁ at the floor −0.92.
v6a and v6b: both channels censored at the floor, ≤ 600 K physical, α₁ at the floor
−0.43 / −0.80. So the E190b/E192 conclusion "the refits moved T_c down" survives the axis
change, and the premise behind the whole v6 programme does not: v5's 734 K sits on FC17's
750 / WAK23's 742 (H_mix-inflection and Landau rows; context, not like-for-like) — the
"rung 1 is ~0.5× published" that v6a/v6b were built to cure was the sampler's factor of
two. The v6 rows (E191/E192 DFT cells) remain valid training data; what they were asked to
fix was not broken. v7 (queued) is the first refit to be judged on the physical axis.
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 12020–12026
E190b — the discriminator. MoNbTaVW at E188's resolution: 6×6×6 (432 sites), 24 states
1800 → 300 K, 100 uncorrelated samples per state (max 1000), annealed. Prediction: if the
sweep was the problem, a peak appears in [550, 950] K; if it still reads ≤ 300 K, reading
(i) holds and rung 1 cannot come from v5 as fitted — the within-composition ranking (ρ 0.80
on RHEA's ordered cells) would have to be strengthened with ordered DFT cells before any
T_c is trusted, and E188's 500 K for Mo–Ta stands only because Mo–Ta's ordering energy is
large enough to survive a weak model. Queued behind E190 (same cores).
EXPERIMENTS.md · lines 12074–12095
E190b result (15:48) — both effects, separated. MoNbTaVW at E188's resolution (432
sites, 24 states 1800 → 300 K, 100 samples/state): T_c = 367 ± 65 K, a clear heat-capacity
peak (Cv 3.64 k_B), NN SRO −0.366 at 300 K — where the 128-site / 60-sample sweep had read
"still rising at 300 K". Published: 750 (FC17), 742 (WAK23). Neither pre-written branch
holds as written: the peak appears (the coarse sweep censored it: resolution effect, real)
but at half the published temperature (model effect, real). Reading, with E188 beside
it: v5's ordering temperatures come out at ~0.5× the published (Mo–Ta 500 vs 600–1000;
MoNbTaVW 367 vs 745; CrTaVW 435 vs 1250 even coarsely) — the within-composition energy
scale is weak by about 2×, while the formation energy scale is right to 10 meV/atom.
The two are different quantities: formation energy is set by the composition-wise mean the
54-atom random cells pin down; the ordering energy is the spread between decorations, which
RHEA's 2–16-atom ordered cells carry with their own rattle-correction noise (12 meV floor
against a 29 meV spread). The icet CE's excess (Mo–Ta 1782) and v5's deficit bracket the
truth from opposite sides.
Consequences: (1) rung 1 stays on hold; (2) any v5 sweep is run at ≥ 432 sites and
≥ 100 samples — the 128-site validation numbers are withdrawn as measurements (kept as
the record of a resolution failure); (3) E188's Mo–Ta 500 K is provisional in value but
robust in sign: a 2× correction puts it at ~1000 K, the published band is 600–1000 — every
estimate except the icet CE's 1782 keeps Mo–Ta's transition inside the window, so the
basin's flip from pass to fail stands; (4) branch → task 23: measure the ordering-energy
scale directly with DFT — random vs B2-type decorations of one composition in 16-atom cells,
minutes each — and feed those cells to v6.
EXPERIMENTS.md · lines 14968–14977
E190b on v5 / v6a / v6b, re-read on the physical axis (2026-09-21 20:2x). The sweeps
were 1800 → 300 K nominal = 3600 → 600 K physical. v5 MoNbTaVW: variance peak 367 K nominal
= 734 K physical, inside the sweep (dE/dT censored at the floor); α₁ at the floor −0.92.
v6a and v6b: both channels censored at the floor, ≤ 600 K physical, α₁ at the floor
−0.43 / −0.80. So the E190b/E192 conclusion "the refits moved T_c down" survives the axis
change, and the premise behind the whole v6 programme does not: v5's 734 K sits on FC17's
750 / WAK23's 742 (H_mix-inflection and Landau rows; context, not like-for-like) — the
"rung 1 is ~0.5× published" that v6a/v6b were built to cure was the sampler's factor of
two. The v6 rows (E191/E192 DFT cells) remain valid training data; what they were asked to
fix was not broken. v7 (queued) is the first refit to be judged on the physical axis.
Related entries
E188 — the ordering temperature from v5, against the published Mo–Ta number
E190 — rung 1 from v5, scored against the published ordering table
E192 — active learning for ordering (2026-09-18 16:30; Liu/Eisenbach 2021 protocol)