Are the ordered states found by the model's own simulation real, checked by quantum calculation?
Partly. For the five-element alloy the model is 25 meV/atom too deep on every such state; for molybdenum–tantalum it is 23 too shallow.
In the log: active learning for ordering (2026-09-18 16:30; Liu/Eisenbach 2021 protocol)
mixedDate 2026-09-18 16:30, as written in the logrung 4 · DFT0 predictions · 4 result paragraphsEXPERIMENTS.md lines 12169–12188, lines 12190–12194, lines 12196–12203, lines 12648–12658, line 12662, lines 12701–12709, lines 12785–12803, line 12805, line 12807, lines 12875–12884
What E192 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E192.svg).
Pre-registration
The pre-registration, as written
E192 — active learning for ordering (2026-09-18 16:30; Liu/Eisenbach 2021 protocol)
Rung 1's deficit (E190/E190b: v5's T_c ≈ 0.5× published even where its formation energies
are right) matches the published diagnosis: a model fitted to random cells and small
ordered compounds has never seen the partially ordered states that set T_c. The remedy
those authors measured is to let the model's own Monte Carlo propose them, compute them in
DFT, and refit. Step 1 (now, CPU, load-gated): v5 canonical MC on 3×3×3 (54 sites) for
Mo₀.₅Ta₀.₅ and MoNbTaVW at 800, 600, 450 and 300 K, saving the final configuration of
each state (--path_to_configs). Step 2 (behind E191): those eight 54-atom cells to DFT
at the standard (~1 h each on 6 ranks). Step 3: E191's eight 16-atom cells + these eight
into training with the QE-source tag → v6, refit; E190b re-run on v6.
Predictions. 1. The sampled configurations are partially ordered: NN Mo–Ta Warren–Cowley
between −0.15 and −0.45 at 300–450 K, near 0 at 800 K (a check that the sampler produced
what the loop needs). 2. v5's own energy on these cells is lower than DFT's by 20–40
meV/atom — the model over-stabilises the ordered states it invents, the signature of a
model extrapolating outside its training distribution — or higher by that much, in which
case it under-orders and T_c is low for that reason; the sign is the finding. 3. v6's T_c
for MoNbTaVW (E190b protocol) rises from 367 to ≥ 500 K; if it does not move by more
than 65 K (E190b's sigma), one round is too few and the loop runs again with the next
MC's states — the paper needed several rounds.
E192 note (16:35) — what the sampler found, before DFT. v5 on its own sampled Mo–Ta cells:
−101.5 (800 K, NN SRO −0.05), −105.5 (633), −127.6 (467), −152.5 (300 K, NN SRO −0.49 — E191's
B2 to the digit; a 3×3×3 conventional cell hosts B2). Eight random decorations of the same
composition on the same cell: −76 ± 5, NN SRO within ±0.07 of zero. So the 800 K state,
random at the first shell, is 25 meV/atom below random under v5: the model rewards an
order beyond nearest neighbours that the sampler exploits. Whether DFT pays for it too is
prediction 2's sign. (A recount of NN order with my own shell window disagreed with pyeCE at
300 K — my window swept in the second shell; pyeCE's numbers are the ones used.)
E192, first cell (2026-09-19 03:38): MoNbTaVW sampled at 300 K. DFT E_form −55.2, v5 −85.7
on the same 54-atom configuration → Δ(v5 − DFT) = −30.5 meV/atom, inside prediction 2's
first band (−40…−20): v5 over-stabilises the ordered state its own MC invents. Set
beside E191: on the guessed quinary ordered cell v5 was 18 meV too shallow; on the state
its sampler chose it is 30 meV too deep. Both are what a model extrapolating outside its
training distribution does — the sampler walks downhill on the model's errors and finds
minima that are not there (Liu/Eisenbach 2021, exactly). This is the row v6 needs most; it
is stored (data/dft/E192_MoNbTaVW_T300) and enters the fit with the site order fixed.
Seven cells follow (466, 633, 800 K for the quinary; four for Mo–Ta). The sign at the hot
end (800 K, near-random) is the control: Δ there should be near 0 if this is extrapolation
and not an offset.
E192, third cell (2026-09-19 10:54): MoNbTaVW sampled at 633 K. DFT −44.7, v5 −74.0 → Δ = −29.4.
The "shrinks with the degree of order" reading from two points is withdrawn: 300/466/633 K
give −30.5 / −19.4 / −29.4, mean −26.4 ± 6.1, no monotone trend. What stands: v5 is
~26 meV too deep on every state its sampler produced for this composition. Two readings
remain and the 800 K cell (near-random SRO by the sampler's own measure) decides between
them: Δ(800 K) near 0 → extrapolation onto ordered states; Δ(800 K) ≈ −26 → a composition-
level offset at MoNbTaVW that has nothing to do with ordering — in which case the v5
absolute energy is wrong there by 26 while ΔE_order may still be right, the Mo–Ta pattern
with the opposite sign. Stored as E192_MoNbTaVW_T633.
E192 Mo–Ta, 300 K (2026-09-19 16:09). The sampler's 300 K state is B2 (MC mean −152.4 = v5's B2 −152.5): DFT −184.5, identical to E191's 16-atom B2 (−184.5) — no cell-size dependence in the DFT frame. Δ(v5 − DFT) = +32.0: v5 too shallow on Mo–Ta's ground state, against −25 too deep on the quinary's sampled states. One model, two systems, two signs of ordering error, both ~30 meV; the deep-end under-fit (training set: +31 on this very row) and the sampler's selection of low-error states are different mechanisms and both are real. 466/633/800 K follow. Row E192_MoTa_T300.
E192 Mo–Ta, 466 K (2026-09-19 19:00). Partially ordered state (MC mean −142.9): DFT −152.5, v5 −127.6 → Δ = +24.9; with B2 (+32.0) the Mo–Ta mean is +28.4 over two. The 633 and 800 K cells decide the reading: Δ → 0 as the state randomises means v5 under-estimates Mo–Ta's ordering stabilisation (the quinary's mirror image); Δ staying ≈ +28 means a flat offset on the sampler's Mo–Ta states. Row E192_MoTa_T466; 6/8 stored.
Results
EXPERIMENTS.md · line 12190
E192 step 1 done (16:34) — prediction 1 confirmed. NN Mo–Ta Warren–Cowley from the
sampler: Mo–Ta −0.05 (800 K), −0.13, −0.35, −0.49 (300 K); MoNbTaVW +0.02, −0.01, −0.10,
−0.34 — near zero hot, partially ordered cold, as required. Eight 54-atom cells built at
Vegard a from the final configurations (v5's own energy on each recorded in
runs/e192_ordering_al/cells.json before DFT); DFT queued behind E191.
EXPERIMENTS.md · line 12662
E192, second cell (2026-09-19 06:49): MoNbTaVW sampled at 466 K. DFT −45.4, v5 −64.8 → Δ = −19.4 (300 K: −30.5). Mean −25.0 ± 7.8 over two. The over-stabilisation shrinks with the degree of order the sampler reached, as extrapolation should and an offset would not; 633 and 800 K complete the curve. Stored as E192_MoNbTaVW_T466.
EXPERIMENTS.md · line 12785
E192 quinary complete (2026-09-19 15:28) — prediction 2, first branch, with its mechanism.
sampled at
DFT
v5 (same cell)
Δ v5−DFT
v5 below its random
DFT below random
300 K
−55.2
−85.7
−30.5
58
28
466 K
−45.4
−64.8
−19.4
38
18
633 K
−44.7
−74.0
−29.4
47
17
800 K
−20.6
−41.4
−20.7
14
−7 (random)
Random reference: v5's own random decoration −27.3 ± 6.4; RHEA's 113 held-out 54-atom
MoNbTaVW random cells −27.8 (v5 predicts them at −30.7, residual −2.9). No frame offset
at this composition — the offset reading is rejected. On the states its sampler chose, v5
is −25.0 ± 5.7 too deep at every temperature, and its ordering stabilisation (energy
below random) is about twice DFT's: 58 vs 28 at 300 K, 14 vs 0 at 800 K. The sampler
selects the configurations on which v5 errs low — Liu & Eisenbach 2021's failure mode,
measured here on the model's own trajectory. These four rows are the ones v6 needs, and
E190b-on-v6b's ≥ 500 K bar is the test that they were enough. Set beside E191's guessed
ordered cell (v5 18 too shallow): v5's error is structure-dependent and the sampler
finds its negative side. Mo–Ta's four sampled cells follow (T300 running).
EXPERIMENTS.md · line 12875
E192 COMPLETE (01:05), 8/8 cells — two systems, two different pathologies.
system
Δ(v5 − DFT) on sampled states
random cells
reading
MoNbTaVW
−25.0 ± 5.7 (−20.7, −29.4, −19.4, −30.5)
agree with DFT to 3 meV
no offset; the sampler selects states where v5 errs low (Liu & Eisenbach)
a composition-level offset at Mo₅₀Ta₅₀, ordered and random alike; ΔE_order survives it (E191: 80.8 vs 80.3)
The eight rows are the v6 training set. v6b's E190b ≥ 500 K bar tests whether they fix the
quinary; the Mo–Ta offset is a separate defect that weighting will not touch — it needs the
Mo–Ta corner in training, which is task 20's DFT hull rebuild.
The full record
This entry is written in 10 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 12169–12188
E192 — active learning for ordering (2026-09-18 16:30; Liu/Eisenbach 2021 protocol)
Rung 1's deficit (E190/E190b: v5's T_c ≈ 0.5× published even where its formation energies
are right) matches the published diagnosis: a model fitted to random cells and small
ordered compounds has never seen the partially ordered states that set T_c. The remedy
those authors measured is to let the model's own Monte Carlo propose them, compute them in
DFT, and refit. Step 1 (now, CPU, load-gated): v5 canonical MC on 3×3×3 (54 sites) for
Mo₀.₅Ta₀.₅ and MoNbTaVW at 800, 600, 450 and 300 K, saving the final configuration of
each state (--path_to_configs). Step 2 (behind E191): those eight 54-atom cells to DFT
at the standard (~1 h each on 6 ranks). Step 3: E191's eight 16-atom cells + these eight
into training with the QE-source tag → v6, refit; E190b re-run on v6.
Predictions. 1. The sampled configurations are partially ordered: NN Mo–Ta Warren–Cowley
between −0.15 and −0.45 at 300–450 K, near 0 at 800 K (a check that the sampler produced
what the loop needs). 2. v5's own energy on these cells is lower than DFT's by 20–40
meV/atom — the model over-stabilises the ordered states it invents, the signature of a
model extrapolating outside its training distribution — or higher by that much, in which
case it under-orders and T_c is low for that reason; the sign is the finding. 3. v6's T_c
for MoNbTaVW (E190b protocol) rises from 367 to ≥ 500 K; if it does not move by more
than 65 K (E190b's sigma), one round is too few and the loop runs again with the next
MC's states — the paper needed several rounds.
EXPERIMENTS.md · lines 12190–12194
E192 step 1 done (16:34) — prediction 1 confirmed. NN Mo–Ta Warren–Cowley from the
sampler: Mo–Ta −0.05 (800 K), −0.13, −0.35, −0.49 (300 K); MoNbTaVW +0.02, −0.01, −0.10,
−0.34 — near zero hot, partially ordered cold, as required. Eight 54-atom cells built at
Vegard a from the final configurations (v5's own energy on each recorded in
runs/e192_ordering_al/cells.json before DFT); DFT queued behind E191.
EXPERIMENTS.md · lines 12196–12203
E192 note (16:35) — what the sampler found, before DFT. v5 on its own sampled Mo–Ta cells:
−101.5 (800 K, NN SRO −0.05), −105.5 (633), −127.6 (467), −152.5 (300 K, NN SRO −0.49 — E191's
B2 to the digit; a 3×3×3 conventional cell hosts B2). Eight random decorations of the same
composition on the same cell: −76 ± 5, NN SRO within ±0.07 of zero. So the 800 K state,
random at the first shell, is 25 meV/atom below random under v5: the model rewards an
order beyond nearest neighbours that the sampler exploits. Whether DFT pays for it too is
prediction 2's sign. (A recount of NN order with my own shell window disagreed with pyeCE at
300 K — my window swept in the second shell; pyeCE's numbers are the ones used.)
EXPERIMENTS.md · lines 12648–12658
E192, first cell (2026-09-19 03:38): MoNbTaVW sampled at 300 K. DFT E_form −55.2, v5 −85.7
on the same 54-atom configuration → Δ(v5 − DFT) = −30.5 meV/atom, inside prediction 2's
first band (−40…−20): v5 over-stabilises the ordered state its own MC invents. Set
beside E191: on the guessed quinary ordered cell v5 was 18 meV too shallow; on the state
its sampler chose it is 30 meV too deep. Both are what a model extrapolating outside its
training distribution does — the sampler walks downhill on the model's errors and finds
minima that are not there (Liu/Eisenbach 2021, exactly). This is the row v6 needs most; it
is stored (data/dft/E192_MoNbTaVW_T300) and enters the fit with the site order fixed.
Seven cells follow (466, 633, 800 K for the quinary; four for Mo–Ta). The sign at the hot
end (800 K, near-random) is the control: Δ there should be near 0 if this is extrapolation
and not an offset.
EXPERIMENTS.md · line 12662
E192, second cell (2026-09-19 06:49): MoNbTaVW sampled at 466 K. DFT −45.4, v5 −64.8 → Δ = −19.4 (300 K: −30.5). Mean −25.0 ± 7.8 over two. The over-stabilisation shrinks with the degree of order the sampler reached, as extrapolation should and an offset would not; 633 and 800 K complete the curve. Stored as E192_MoNbTaVW_T466.
EXPERIMENTS.md · lines 12701–12709
E192, third cell (2026-09-19 10:54): MoNbTaVW sampled at 633 K. DFT −44.7, v5 −74.0 → Δ = −29.4.
The "shrinks with the degree of order" reading from two points is withdrawn: 300/466/633 K
give −30.5 / −19.4 / −29.4, mean −26.4 ± 6.1, no monotone trend. What stands: v5 is
~26 meV too deep on every state its sampler produced for this composition. Two readings
remain and the 800 K cell (near-random SRO by the sampler's own measure) decides between
them: Δ(800 K) near 0 → extrapolation onto ordered states; Δ(800 K) ≈ −26 → a composition-
level offset at MoNbTaVW that has nothing to do with ordering — in which case the v5
absolute energy is wrong there by 26 while ΔE_order may still be right, the Mo–Ta pattern
with the opposite sign. Stored as E192_MoNbTaVW_T633.
EXPERIMENTS.md · lines 12785–12803
E192 quinary complete (2026-09-19 15:28) — prediction 2, first branch, with its mechanism.
sampled at
DFT
v5 (same cell)
Δ v5−DFT
v5 below its random
DFT below random
300 K
−55.2
−85.7
−30.5
58
28
466 K
−45.4
−64.8
−19.4
38
18
633 K
−44.7
−74.0
−29.4
47
17
800 K
−20.6
−41.4
−20.7
14
−7 (random)
Random reference: v5's own random decoration −27.3 ± 6.4; RHEA's 113 held-out 54-atom
MoNbTaVW random cells −27.8 (v5 predicts them at −30.7, residual −2.9). No frame offset
at this composition — the offset reading is rejected. On the states its sampler chose, v5
is −25.0 ± 5.7 too deep at every temperature, and its ordering stabilisation (energy
below random) is about twice DFT's: 58 vs 28 at 300 K, 14 vs 0 at 800 K. The sampler
selects the configurations on which v5 errs low — Liu & Eisenbach 2021's failure mode,
measured here on the model's own trajectory. These four rows are the ones v6 needs, and
E190b-on-v6b's ≥ 500 K bar is the test that they were enough. Set beside E191's guessed
ordered cell (v5 18 too shallow): v5's error is structure-dependent and the sampler
finds its negative side. Mo–Ta's four sampled cells follow (T300 running).
EXPERIMENTS.md · line 12805
E192 Mo–Ta, 300 K (2026-09-19 16:09). The sampler's 300 K state is B2 (MC mean −152.4 = v5's B2 −152.5): DFT −184.5, identical to E191's 16-atom B2 (−184.5) — no cell-size dependence in the DFT frame. Δ(v5 − DFT) = +32.0: v5 too shallow on Mo–Ta's ground state, against −25 too deep on the quinary's sampled states. One model, two systems, two signs of ordering error, both ~30 meV; the deep-end under-fit (training set: +31 on this very row) and the sampler's selection of low-error states are different mechanisms and both are real. 466/633/800 K follow. Row E192_MoTa_T300.
EXPERIMENTS.md · line 12807
E192 Mo–Ta, 466 K (2026-09-19 19:00). Partially ordered state (MC mean −142.9): DFT −152.5, v5 −127.6 → Δ = +24.9; with B2 (+32.0) the Mo–Ta mean is +28.4 over two. The 633 and 800 K cells decide the reading: Δ → 0 as the state randomises means v5 under-estimates Mo–Ta's ordering stabilisation (the quinary's mirror image); Δ staying ≈ +28 means a flat offset on the sampler's Mo–Ta states. Row E192_MoTa_T466; 6/8 stored.
EXPERIMENTS.md · lines 12875–12884
E192 COMPLETE (01:05), 8/8 cells — two systems, two different pathologies.
system
Δ(v5 − DFT) on sampled states
random cells
reading
MoNbTaVW
−25.0 ± 5.7 (−20.7, −29.4, −19.4, −30.5)
agree with DFT to 3 meV
no offset; the sampler selects states where v5 errs low (Liu & Eisenbach)
a composition-level offset at Mo₅₀Ta₅₀, ordered and random alike; ΔE_order survives it (E191: 80.8 vs 80.3)
The eight rows are the v6 training set. v6b's E190b ≥ 500 K bar tests whether they fix the
quinary; the Mo–Ta offset is a separate defect that weighting will not touch — it needs the
Mo–Ta corner in training, which is task 20's DFT hull rebuild.
Related entries
E190 — rung 1 from v5, scored against the published ordering table