Experiments · E191

Does the energy model get the energy released by ordering right, checked by quantum calculation?

Partly. For molybdenum–tantalum yes, 80.3 against 80.8 meV/atom; for the five-element alloy the ordered cell was a guess, so no ratio stands.

In the log: the ordering-energy scale, measured directly (15:49, cells built, DFT queued)

mixedDate 2026-09-19 00:00, as written in the logrung 4 · DFT2 predictions · 0 result paragraphsEXPERIMENTS.md lines 12097–12116, lines 12335–12356, lines 12358–12362, lines 12409–12438, lines 12440–12445, lines 12447–12493, lines 12529–12538, lines 12572–12582, lines 13441–13445, lines 13656–13719
exp E191 diagram
What E191 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E191.svg).

Pre-registration

  1. (2)
    , that the deep-row residual moves by less than 3 meV, would have been scored against a corrupted set. Fixed, in this order: the row was moved out of data/dft to runs/quarantine_wrong_lattice/; the chain and the in-flight pw.x (MoNbTaVW, already running at 3.2935) were stopped; ground_state.py now rescales to BCC.cell_length(comp) before writing cell.vasp/pw.in and prints both constants; all three cells were regenerated (Mo–Ta 3.2460, MoNbTaVW 3.2049, MoNbTaW 3.2556) and re-verified as intact B2 and pseudo-binary B2; DFT relaunched on the corrected cells. Prediction
    no verdict written against it
  2. (4)
    still has to run.
    no verdict written against it
The pre-registration, as written

E191 — the ordering-energy scale, measured directly (15:49, cells built, DFT queued)

16-atom (2×2×2 conventional) cells at the settings standard (60/720 Ry, MV 0.02, k 7×7×7, Vegard a): for Mo₀.₅Ta₀.₅ the exact B2 (Mo vertices, Ta centres) and three random decorations; for MoNbTaVW (3/3/3/3 + 4 V in 16) a Widom-type ordering (Mo, W + 2 V on centres; Nb, Ta + 2 V on vertices) and three randoms. ΔE_order = ⟨E_random⟩ − E_ordered on the same cells in both DFT and v5, so no reference enters. v5, before DFT: Mo–Ta: B2 −152.5, random −72.2 ± 15.8 → ΔE_order(v5) = +80.3 meV/atom (literature: B2 −186 (Widom, PBE) against random −90 to −127 → 60–100). MoNbTaVW: ordered −14.0, random −10.5 ± 19.2 → +3.6 — the three-cell spread exceeds the mean, and the ordered decoration is a guess; this half is a weak instrument and is scored as such. Predictions. 1. As registered in task 23: DFT ΔE_order(Mo–Ta) in [115, 200] (v5 at 0.4–0.7×). 2. The alternative, stated now because v5's 80 sits in the literature range: DFT in [70, 95] (v5 at 0.85–1.15×) — then the ordering energy is right and the T_c deficit (E188 500 vs 600–1000; E190b 367 vs 745) lives in the MC/entropy side (sample count, cell size, or a many-body term the sweep under-resolves), not the model's scale, and rung 1's fix is the sweep protocol, not the fit. 3. Below 70: v5 over-orders Mo–Ta and the T_c deficit is elsewhere entirely. 4. Quinary: DFT ΔE_order ≥ 15 meV (the published 745 K needs a real ordering energy); v5's 3.6 is then a 4×+ under-read of a weak ordering — the composition where the model is least anchored.

E191, partial (23:15; MoNbTaVW B2 + 2 of 3 randoms done, Mo–Ta not started). MoNbTaVW: ΔE_order(DFT) = ⟨E_random⟩ − E_B2 = +21.8 meV/atom (random sd 2.5 over two cells) against v5's +3.6 → v5/DFT = 0.16. Prediction 4 (≥ 15 meV) confirmed on partial data: the published 745 K rests on a real ordering energy, and v5 reads it six times too small at the composition where it is least anchored. Final with rand2, then Mo–Ta (the [115, 200] / [70, 95] / < 70 branches). Verdict scripts written and tested ahead of the collection: scripts/ordering/e191_verdict.py, scripts/ordering/e192_verdict.py (pending cells listed, never skipped; ΔE_order from raw energies, no reference enters).

N4 queued as a chain (23:18): E191 + E192 → store → v6 → E190b on v6. Recipe fixed now so the wake-up collects rather than improvises. v5's recipe, recovered from its counts (4167 trained + 159 held = every label): train_ece.py defaults on the pureref9 labels, no trusted-only filter. v6 differs in one flag: --holdout "Mo,Nb,Ta,W". The default also holds out every {Mo,Nb,Ta,V,W} row by element set, which would divert E191's and E192's MoNbTaVW cells — the rows v6 exists to learn from — into the test file; MoNbTaW stays held out so the T=0 ordered-vs-random check keeps an untouched system. Store rows come from scripts/expansion/store_qe_cell.py (settings read from the input that ran, refuses unfinished or duplicate runs). Caveat on record: E191's 16-atom cells are at 7×7×7 (spacing 0.138 Å⁻¹) against the references' 0.152 and the 54-atom cells' 0.163; ΔE_order is immune, E_form may carry a few meV of mesh convention. E190b-on-v6 prediction stands: ≥ 500 K (E192 prediction 3), decision bar ±65 K. Chain: runs/v6_chain.sh (waits for E192's DFT and for both running searches to end).

E191 quinary cells stored, per-cell v5 − DFT (23:19; 16-atom, 7×7×7, refs_v3_fit). MoNbTaVW B2: DFT −35.0, v5 −14.0 → v5 +21.0 too high on the ordered cell; rand0: DFT −15.0, v5 −20.3 (−5.3); rand1: DFT −11.5, v5 +16.3 (+27.8). The ordered state is what v5 misses, in the direction that lowers T_c (under-ordering), and the random cells scatter by ±30 — a 16-atom random cell is one configuration, not an average. Rows in data/dft/E191_*.

E191 quinary — final (23:30; B2 + 3 randoms, 16 atoms, 60/720, 7×7×7). MoNbTaVW: E_B2 = −35.0, ⟨E_random⟩ = −13.4 ± 1.8 meV/atom (−15.0 / −11.5 / −13.8) → ΔE_order(DFT) = +21.6 ± 1.8; v5 gives +3.6 → v5/DFT = 0.16. Prediction 4 confirmed: the published 745 K rests on a real ordering energy and v5 reads it at one sixth. The three random cells agree to 1.8 meV, so the number is the model's, not the sampling's. Mo–Ta (branches [115, 200] / [70, 95] / < 70) follows; its B2 cell is running.

N6 — measured cause (23:31): the FB site sits one step beyond the propagation frontier. E171b head rebuilt in-process (steps 4, broadcast, LAL+DNa+MBON+PFL+DN readout, seed 2, CPU): fan-shaped-body cells are all active at the cached step (eligibility nonzero on 100 % of the site's edges, max 1.3e-2), four FB lessons move the FB→hΔ/vΔ weights by |Δw| = 27.9 (max 0.026 per edge) — and the head's scores change by exactly 0.0. The 5-HT control (LAL→DN) under the same lessons moves its weights by 1.6 and the scores by 2.7e-8: nonzero. Reading: FB is first reached at step 4 (E171 measured "4 reaches every central-complex cell"), so edges leaving FB are traversed at step 5, which a four-step forward never runs. Every FB lesson in E166/E171b/E173 rewrote weights the readout never sees. E171b ≡ E178 is therefore exact, not statistical. The step count the FB site needs (5 or 6: FB→hΔ→ PFL→readout is two hops past the frontier) is being measured now; E199 = fb,oa,5ht vs oa,5ht at that depth, same code, --confirm, ≥ 2 seeds. Steps > 4 changes every node's state, so E199 needs its own no-FB control at the same depth — the E171b/E178 lesson, applied.

FB depth measured (23:37) and E199 queued. Same diagnostic at steps 5 and 6: four FB lessons now move the scores by 9.0e-9 and 9.6e-9 (5-HT control 4.5e-8 / 4.0e-8), against exactly 0 at steps 4. The site acts from step 5 on. E199 — fb,oa,5ht vs oa,5ht, both at FORAGER_STEPS=5, --confirm, FB lr 2.0, eight-element space on the v5 reward, MPS, 2 seeds each, one code version. Prediction, on record: null — |AUC_Q(fb) − AUC_Q(no fb)| within one seed-sd; the per-lesson score movement is 2e-5 of the score, and the 5-HT site, which moves scores by five times more per lesson, is itself undecided (E173 vs E173c). If FB clears one seed-sd the site carries something and the book's FB chapter is rewritten on E199, not E171b. Chain runs/e199_chain.sh (after E198; v6 waits for it).

E191 Mo–Ta B2 landed (23:38). E_form(DFT) = −184.5 meV/atom (16-atom B2, 60/720, 7×7×7, refs_v3_fit). Widom's PBE B2 MoTa is −186: the DFT standard reproduces the one published number in this experiment to 2 meV. v5 on the same cell: −152.5 → v5 is 32.0 meV too high on the ordered state, the same sign as the quinary (+21.0 on its B2). The three Mo–Ta random cells decide the branch; literature random −90 to −127 would put ΔE_order(DFT) at 60–95, i.e. prediction 2's band or just above it. Row data/dft/E191_MoTa_B2.

E191 Mo–Ta, partial (23:46; B2 + rand0). rand0: DFT −123.1, v5 −91.3 on the same 16-atom configuration → v5 +31.8 too high; B2 was +32.0. Partial ΔE_order(DFT) = +61.4 vs v5 +80.3 (v5/DFT 1.31) — the partial branch is prediction 3 ("v5 over-orders Mo–Ta"), but with one random cell it is not a branch yet. What the two cells do say: at Mo50Ta50 in 16-atom cells v5 sits ~32 meV above DFT on ordered and random alike, whereas on 54-atom Mo–Ta cells (E180, and the RHEA rows it trained on) it agrees within 10 meV. Hypothesis, to be tested by E192's 54-atom cells and the v6 fit: the 6 Å neighbourhood in a 6.5 Å cell contains an atom's own periodic image, an environment the 54-atom training set never shows the embedding — a cell-size bias, not an ordering bias. If E192's 54-atom cells show no such offset, E191's absolute energies carry it and only ΔE_order (offset-free) is used.

Convention closed, hypothesis withdrawn, and a sharper cause (23:48). (i) The cell-size hypothesis above is withdrawn: 2,611 of v5's 4,326 training rows are 16-atom cells and its held-out residual on 16-atom cells is +2.9 mean / 6.9 MAE (54-atom: −3.0 / 13.4). (ii) RHEA's own 16-atom Mo₈Ta₈ ordered row (bcc_alloys_ordered) is −183.9; E191's B2 at the QE standard is −184.5 — VASP/RHEA-references and QE/refs_v3_fit agree to 0.6 meV, so E191's absolute energies are on the training set's convention. (iii) Hence v5's −152.5 on B2 MoTa is a 31 meV under-fit of a row in its own training set, the deepest Mo–Ta point. That is a fit-quality statement, and if it holds across the deep ordered rows it is the rung-1 mechanism: the ordering energy is small because the model regresses the deepest ordered states toward the mean — which adding a few E192 rows to v6 will not cure without weighting or capacity. Measured next on v5's own training rows.

Rung-1 mechanism, measured on v5's own training set (23:49). v5 re-predicted on its training structures: deepest 60 rows (label mean −104.1, min −183.9): residual v5 − label +10.0 mean, MAE 18.3, worst +61.9; 60 random rows (label mean +64.3): +0.7 mean, MAE 9.8. Deepest ten: B2 MoTa +31.3, three rows at −142/−133 each +35, the rest within 5. The model regresses the deep ordered states toward the mean — one-signed, +10 on average and +30–35 on the deepest — so every ordering energy it computes as E_ordered − E_random is too small by that amount, and T_c with it. This is the rung-1 deficit's origin, measured; the RHEA rows that carry ordering were in the fit all along. Consequence for N4: v6 on the same recipe inherits the regression; the refit needs the deep rows weighted (or capacity), tested as a pair: v6a (plain, queued) vs v6b (rows below −100 meV weighted ×4), both scored on E190b (≥ 500 K bar) and on this deep-60 residual.

Flywheel bug caught before v6 (23:50). train_ece.py's emitter writes every row's symbols against ideal_frac_of(n), a fixed site order — right for RHEA rows, wrong for QE store rows, whose symbols are in file order: both stored cells checked (E191 B2, E187) have the ideal site set but not that order, so v6 would have trained on scrambled configurations — E191's B2 as a random cell. No fit has consumed a QE row yet. Fix: symbols_on_ideal_sites re-places a row's symbols by position (refuses non-ideal sites); test with a permuted B2. Added --deep-weight/--deep-below (row_weight) for v6b. N4 is now a pair on one chain: v6a (plain) and v6b (rows below −100 meV weighted ×4), each followed by E190b (MoNbTaVW, 6³ cells, 24 temperatures, 100 samples). Predictions: v6a T_c within 65 K of v5's 367 (the regression is untouched); v6b ≥ 500 K and its deep-60 residual mean within ±3 meV (from +10). If v6b clears the bar the rung-1 fix is weighting, not data; if neither moves, capacity (embedding/layers) is next.

E191 Mo–Ta, 2 of 3 randoms (2026-09-19 00:00). rand1: DFT −92.4, v5 −52.7 (+39.7). Three cells, three one-signed misses: B2 +32.0, rand0 +31.8, rand1 +39.7 — v5 is ~35 meV too shallow across Mo₅₀Ta₅₀ 16-atom cells regardless of order. Yet the difference is right: ΔE_order(DFT) = +76.8 (random sd 21.7) vs v5 +80.3, ratio 1.05 — partial branch = prediction 2: the Mo–Ta ordering energy is correct and the E188 deficit (500 K vs 600–1000) is on the MC/entropy side. The quinary said the opposite (ratio 0.16). Two systems, two mechanisms: at MoNbTaVW the model's ordering scale is wrong; at Mo–Ta the scale is right and something in the sweep protocol (cell 6³, 100 samples, 24 T-steps, or the canonical ensemble's finite-size rounding) censors the peak. rand2 decides the branch; then E190b-on-v6b must be read per system, not as one number.

E191 FINAL (2026-09-19 00:12) — N3 closed. Mo–Ta: E_B2 −184.5; randoms −123.1 / −92.4 / −95.7 (mean −103.7 ± 16.8) → ΔE_order(DFT) = +80.8 ± 16.8 vs v5 +80.3 — ratio 0.99. Branch: prediction 2, cleanly: v5's Mo–Ta ordering energy is right to the last meV, and its 500 K (E188) against the published 600–1000 K is not the model's ordering scale. The per-cell offsets are one-signed and large — v5 − DFT = +32.0 (B2), +31.8, +39.7, +23.2 (mean +31.7) — so v5 is uniformly ~32 meV too shallow at Mo₅₀Ta₅₀ in 16-atom cells, an absolute-energy error that cancels in ΔE_order and would not cancel in a formation-energy reward or a hull. Quinary: ΔE_order(DFT) +21.6 ± 1.8 vs v5 +3.6 on the same guessed cell (prediction 4 confirmed as a lower bound; the ratio is withdrawn, see the caveat above). Eight rows in data/dft/E191_*; they enter v6 with the site order fixed. What E200 now tests is the MC side for Mo–Ta; what E192 tests is the quinary on v5's own sampled states.

E191's quinary reference was off by a factor of twenty. ΔE_order(v5) from E191's guessed ordered cell was +3.6; from the model's actual ground state it is +73.7. The "one sixth of DFT" reading was already withdrawn as a guess; this says how far off the guess was. E192's DFT verdict is unaffected — it was measured on MC-sampled cells, not on this one — and prediction (4) still has to run.

E191's B2 Mo–Ta is −184.5 and E192's 54-atom Mo–Ta at the same 4×4×4 mesh is −184.49. Same crystal, same composition, same settings — so a 41 meV gap could not be cell size or k-mesh. It is the lattice constant: ground_state.py wrote cell.vasp straight from to_structure, which inherits pyeCE's PRIM constant a = 3.2935 Å. That is a mapping lattice for the cluster expansion, not a geometry. Every row in data/dft carries the convention "ideal bcc sites, Vegard a", and Vegard for Mo₅₀Ta₅₀ is 3.2460 Å. A 1.46 % expansion, and 41 meV/atom of it.

What this would have cost if it had gone unnoticed. The row was already written into data/dft/, and the v7 chain merges that whole directory into the training labels. A cell 41 meV off, labelled with a convention it does not have, would have gone into the first refit that was supposed to close the flywheel on the deployed model — and v7's own prediction (2), that the deep-row residual moves by less than 3 meV, would have been scored against a corrupted set.

Fixed, in this order: the row was moved out of data/dft to runs/quarantine_wrong_lattice/; the chain and the in-flight pw.x (MoNbTaVW, already running at 3.2935) were stopped; ground_state.py now rescales to BCC.cell_length(comp) before writing cell.vasp/pw.in and prints both constants; all three cells were regenerated (Mo–Ta 3.2460, MoNbTaVW 3.2049, MoNbTaW 3.2556) and re-verified as intact B2 and pseudo-binary B2; DFT relaunched on the corrected cells.

Prediction (4) is not scored yet. Nothing from the first DFT attempt counts. The −143.3 is retained only as the measurement of this defect.

Why it was caught. Not by a test, and not by the run itself — by asking why a number disagreed with a number for the same crystal already in the record. The same habit caught E213's scrambled site order an hour earlier. The v5 search results are unaffected: those are lattice-model evaluations keyed to the PRIM, and the ground states, spreads and ΔE_order values in E211 all stand.

And the defect is older than E211 — two rows already in the store have it (2026-09-21 08:3x). The new guard, run over every existing row, refuses E172_HfMoNbTaTiVWZr_s1.00 and E172b_HfMoNbTaTiVWZr_std: both sit at 3.2935 Å against a linear-Vegard 3.3426 for their composition, 1.47 % compressed, and both are stamped "ideal bcc sites, Vegard a".

But their energies are sound, and that is the more interesting finding. E172 ran a three-point volume scan on exactly that cell and recorded the minimum at a₀ = 3.2919 Å, with E(3.2935) − E(min) = +0.02 meV/atom. So the cell sits essentially on its own energy minimum; it is linear Vegard that is wrong here, by 1.5 %, on a composition spanning V (3.001 Å) to Zr (3.689 Å). Vegard's law is a proxy and large size mismatch breaks it.

So the rule is not "always Vegard", it is "no unstated mismatch". The guard now takes --allow-off-vegard <reason> and every stored row records a_conv_A, a_vegard_A and the reason, so no future row has to be trusted on its stamp. E172/E172b keep their energies and need only a corrected annotation; E211's Mo–Ta cell is a different case entirely — 3.2935 is not near its minimum (Mo–Ta's is at its Vegard 3.2460, confirmed by E191 matching RHEA's VASP row to 0.6 meV and Widom's −186), so its 41 meV really was an error.

The distinction that matters for the flywheel: a rung-4 cell must sit at its own energy minimum, because that is the geometry the eCE's labels correspond to. Linear Vegard is the cheap proxy for it, validated on Mo–Ta and falsified on the Hf/Zr-rich octonary. Any composition with that kind of size mismatch needs a volume scan before its DFT label is trusted — and there is no record of one for any cell other than E172's.

A second-order trap in the same repair (2026-09-21 08:4x). The relaunched DFT chain skipped Mo–Ta. Its guard is grep -q "JOB DONE" pw.out && continue, and the stale 3.2935 Å output still sat there carrying JOB DONE — the earlier attempt to move it had been part of a command that was refused, and I did not re-check before relaunching. So the one cell I most needed re-run was the one the chain declined to touch, and its old pw.out would have been read as the new result. Caught by noticing the "new" run reported a byte-identical total energy (−29458.98113510 Ry). The file is quarantined and runs/e211_mota_rerun.sh is queued behind the other two. A resume guard keyed on a file the repair was supposed to delete will silently preserve exactly what the repair was for.

Results

No result paragraph for this entry was found in the log.

The full record

This entry is written in 10 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 12097–12116

E191 — the ordering-energy scale, measured directly (15:49, cells built, DFT queued)

16-atom (2×2×2 conventional) cells at the settings standard (60/720 Ry, MV 0.02, k 7×7×7, Vegard a): for Mo₀.₅Ta₀.₅ the exact B2 (Mo vertices, Ta centres) and three random decorations; for MoNbTaVW (3/3/3/3 + 4 V in 16) a Widom-type ordering (Mo, W + 2 V on centres; Nb, Ta + 2 V on vertices) and three randoms. ΔE_order = ⟨E_random⟩ − E_ordered on the same cells in both DFT and v5, so no reference enters. v5, before DFT: Mo–Ta: B2 −152.5, random −72.2 ± 15.8 → ΔE_order(v5) = +80.3 meV/atom (literature: B2 −186 (Widom, PBE) against random −90 to −127 → 60–100). MoNbTaVW: ordered −14.0, random −10.5 ± 19.2 → +3.6 — the three-cell spread exceeds the mean, and the ordered decoration is a guess; this half is a weak instrument and is scored as such. Predictions. 1. As registered in task 23: DFT ΔE_order(Mo–Ta) in [115, 200] (v5 at 0.4–0.7×). 2. The alternative, stated now because v5's 80 sits in the literature range: DFT in [70, 95] (v5 at 0.85–1.15×) — then the ordering energy is right and the T_c deficit (E188 500 vs 600–1000; E190b 367 vs 745) lives in the MC/entropy side (sample count, cell size, or a many-body term the sweep under-resolves), not the model's scale, and rung 1's fix is the sweep protocol, not the fit. 3. Below 70: v5 over-orders Mo–Ta and the T_c deficit is elsewhere entirely. 4. Quinary: DFT ΔE_order ≥ 15 meV (the published 745 K needs a real ordering energy); v5's 3.6 is then a 4×+ under-read of a weak ordering — the composition where the model is least anchored.

EXPERIMENTS.md · lines 12335–12356

E191, partial (23:15; MoNbTaVW B2 + 2 of 3 randoms done, Mo–Ta not started). MoNbTaVW: ΔE_order(DFT) = ⟨E_random⟩ − E_B2 = +21.8 meV/atom (random sd 2.5 over two cells) against v5's +3.6 → v5/DFT = 0.16. Prediction 4 (≥ 15 meV) confirmed on partial data: the published 745 K rests on a real ordering energy, and v5 reads it six times too small at the composition where it is least anchored. Final with rand2, then Mo–Ta (the [115, 200] / [70, 95] / < 70 branches). Verdict scripts written and tested ahead of the collection: scripts/ordering/e191_verdict.py, scripts/ordering/e192_verdict.py (pending cells listed, never skipped; ΔE_order from raw energies, no reference enters).

N4 queued as a chain (23:18): E191 + E192 → store → v6 → E190b on v6. Recipe fixed now so the wake-up collects rather than improvises. v5's recipe, recovered from its counts (4167 trained + 159 held = every label): train_ece.py defaults on the pureref9 labels, no trusted-only filter. v6 differs in one flag: --holdout "Mo,Nb,Ta,W". The default also holds out every {Mo,Nb,Ta,V,W} row by element set, which would divert E191's and E192's MoNbTaVW cells — the rows v6 exists to learn from — into the test file; MoNbTaW stays held out so the T=0 ordered-vs-random check keeps an untouched system. Store rows come from scripts/expansion/store_qe_cell.py (settings read from the input that ran, refuses unfinished or duplicate runs). Caveat on record: E191's 16-atom cells are at 7×7×7 (spacing 0.138 Å⁻¹) against the references' 0.152 and the 54-atom cells' 0.163; ΔE_order is immune, E_form may carry a few meV of mesh convention. E190b-on-v6 prediction stands: ≥ 500 K (E192 prediction 3), decision bar ±65 K. Chain: runs/v6_chain.sh (waits for E192's DFT and for both running searches to end).

EXPERIMENTS.md · lines 12358–12362

E191 quinary cells stored, per-cell v5 − DFT (23:19; 16-atom, 7×7×7, refs_v3_fit). MoNbTaVW B2: DFT −35.0, v5 −14.0 → v5 +21.0 too high on the ordered cell; rand0: DFT −15.0, v5 −20.3 (−5.3); rand1: DFT −11.5, v5 +16.3 (+27.8). The ordered state is what v5 misses, in the direction that lowers T_c (under-ordering), and the random cells scatter by ±30 — a 16-atom random cell is one configuration, not an average. Rows in data/dft/E191_*.

EXPERIMENTS.md · lines 12409–12438

E191 quinary — final (23:30; B2 + 3 randoms, 16 atoms, 60/720, 7×7×7). MoNbTaVW: E_B2 = −35.0, ⟨E_random⟩ = −13.4 ± 1.8 meV/atom (−15.0 / −11.5 / −13.8) → ΔE_order(DFT) = +21.6 ± 1.8; v5 gives +3.6 → v5/DFT = 0.16. Prediction 4 confirmed: the published 745 K rests on a real ordering energy and v5 reads it at one sixth. The three random cells agree to 1.8 meV, so the number is the model's, not the sampling's. Mo–Ta (branches [115, 200] / [70, 95] / < 70) follows; its B2 cell is running.

N6 — measured cause (23:31): the FB site sits one step beyond the propagation frontier. E171b head rebuilt in-process (steps 4, broadcast, LAL+DNa+MBON+PFL+DN readout, seed 2, CPU): fan-shaped-body cells are all active at the cached step (eligibility nonzero on 100 % of the site's edges, max 1.3e-2), four FB lessons move the FB→hΔ/vΔ weights by |Δw| = 27.9 (max 0.026 per edge) — and the head's scores change by exactly 0.0. The 5-HT control (LAL→DN) under the same lessons moves its weights by 1.6 and the scores by 2.7e-8: nonzero. Reading: FB is first reached at step 4 (E171 measured "4 reaches every central-complex cell"), so edges leaving FB are traversed at step 5, which a four-step forward never runs. Every FB lesson in E166/E171b/E173 rewrote weights the readout never sees. E171b ≡ E178 is therefore exact, not statistical. The step count the FB site needs (5 or 6: FB→hΔ→ PFL→readout is two hops past the frontier) is being measured now; E199 = fb,oa,5ht vs oa,5ht at that depth, same code, --confirm, ≥ 2 seeds. Steps > 4 changes every node's state, so E199 needs its own no-FB control at the same depth — the E171b/E178 lesson, applied.

FB depth measured (23:37) and E199 queued. Same diagnostic at steps 5 and 6: four FB lessons now move the scores by 9.0e-9 and 9.6e-9 (5-HT control 4.5e-8 / 4.0e-8), against exactly 0 at steps 4. The site acts from step 5 on. E199 — fb,oa,5ht vs oa,5ht, both at FORAGER_STEPS=5, --confirm, FB lr 2.0, eight-element space on the v5 reward, MPS, 2 seeds each, one code version. Prediction, on record: null — |AUC_Q(fb) − AUC_Q(no fb)| within one seed-sd; the per-lesson score movement is 2e-5 of the score, and the 5-HT site, which moves scores by five times more per lesson, is itself undecided (E173 vs E173c). If FB clears one seed-sd the site carries something and the book's FB chapter is rewritten on E199, not E171b. Chain runs/e199_chain.sh (after E198; v6 waits for it).

EXPERIMENTS.md · lines 12440–12445

E191 Mo–Ta B2 landed (23:38). E_form(DFT) = −184.5 meV/atom (16-atom B2, 60/720, 7×7×7, refs_v3_fit). Widom's PBE B2 MoTa is −186: the DFT standard reproduces the one published number in this experiment to 2 meV. v5 on the same cell: −152.5 → v5 is 32.0 meV too high on the ordered state, the same sign as the quinary (+21.0 on its B2). The three Mo–Ta random cells decide the branch; literature random −90 to −127 would put ΔE_order(DFT) at 60–95, i.e. prediction 2's band or just above it. Row data/dft/E191_MoTa_B2.

EXPERIMENTS.md · lines 12447–12493

E191 Mo–Ta, partial (23:46; B2 + rand0). rand0: DFT −123.1, v5 −91.3 on the same 16-atom configuration → v5 +31.8 too high; B2 was +32.0. Partial ΔE_order(DFT) = +61.4 vs v5 +80.3 (v5/DFT 1.31) — the partial branch is prediction 3 ("v5 over-orders Mo–Ta"), but with one random cell it is not a branch yet. What the two cells do say: at Mo50Ta50 in 16-atom cells v5 sits ~32 meV above DFT on ordered and random alike, whereas on 54-atom Mo–Ta cells (E180, and the RHEA rows it trained on) it agrees within 10 meV. Hypothesis, to be tested by E192's 54-atom cells and the v6 fit: the 6 Å neighbourhood in a 6.5 Å cell contains an atom's own periodic image, an environment the 54-atom training set never shows the embedding — a cell-size bias, not an ordering bias. If E192's 54-atom cells show no such offset, E191's absolute energies carry it and only ΔE_order (offset-free) is used.

Convention closed, hypothesis withdrawn, and a sharper cause (23:48). (i) The cell-size hypothesis above is withdrawn: 2,611 of v5's 4,326 training rows are 16-atom cells and its held-out residual on 16-atom cells is +2.9 mean / 6.9 MAE (54-atom: −3.0 / 13.4). (ii) RHEA's own 16-atom Mo₈Ta₈ ordered row (bcc_alloys_ordered) is −183.9; E191's B2 at the QE standard is −184.5 — VASP/RHEA-references and QE/refs_v3_fit agree to 0.6 meV, so E191's absolute energies are on the training set's convention. (iii) Hence v5's −152.5 on B2 MoTa is a 31 meV under-fit of a row in its own training set, the deepest Mo–Ta point. That is a fit-quality statement, and if it holds across the deep ordered rows it is the rung-1 mechanism: the ordering energy is small because the model regresses the deepest ordered states toward the mean — which adding a few E192 rows to v6 will not cure without weighting or capacity. Measured next on v5's own training rows.

Rung-1 mechanism, measured on v5's own training set (23:49). v5 re-predicted on its training structures: deepest 60 rows (label mean −104.1, min −183.9): residual v5 − label +10.0 mean, MAE 18.3, worst +61.9; 60 random rows (label mean +64.3): +0.7 mean, MAE 9.8. Deepest ten: B2 MoTa +31.3, three rows at −142/−133 each +35, the rest within 5. The model regresses the deep ordered states toward the mean — one-signed, +10 on average and +30–35 on the deepest — so every ordering energy it computes as E_ordered − E_random is too small by that amount, and T_c with it. This is the rung-1 deficit's origin, measured; the RHEA rows that carry ordering were in the fit all along. Consequence for N4: v6 on the same recipe inherits the regression; the refit needs the deep rows weighted (or capacity), tested as a pair: v6a (plain, queued) vs v6b (rows below −100 meV weighted ×4), both scored on E190b (≥ 500 K bar) and on this deep-60 residual.

Flywheel bug caught before v6 (23:50). train_ece.py's emitter writes every row's symbols against ideal_frac_of(n), a fixed site order — right for RHEA rows, wrong for QE store rows, whose symbols are in file order: both stored cells checked (E191 B2, E187) have the ideal site set but not that order, so v6 would have trained on scrambled configurations — E191's B2 as a random cell. No fit has consumed a QE row yet. Fix: symbols_on_ideal_sites re-places a row's symbols by position (refuses non-ideal sites); test with a permuted B2. Added --deep-weight/--deep-below (row_weight) for v6b. N4 is now a pair on one chain: v6a (plain) and v6b (rows below −100 meV weighted ×4), each followed by E190b (MoNbTaVW, 6³ cells, 24 temperatures, 100 samples). Predictions: v6a T_c within 65 K of v5's 367 (the regression is untouched); v6b ≥ 500 K and its deep-60 residual mean within ±3 meV (from +10). If v6b clears the bar the rung-1 fix is weighting, not data; if neither moves, capacity (embedding/layers) is next.

EXPERIMENTS.md · lines 12529–12538

E191 Mo–Ta, 2 of 3 randoms (2026-09-19 00:00). rand1: DFT −92.4, v5 −52.7 (+39.7). Three cells, three one-signed misses: B2 +32.0, rand0 +31.8, rand1 +39.7 — v5 is ~35 meV too shallow across Mo₅₀Ta₅₀ 16-atom cells regardless of order. Yet the difference is right: ΔE_order(DFT) = +76.8 (random sd 21.7) vs v5 +80.3, ratio 1.05 — partial branch = prediction 2: the Mo–Ta ordering energy is correct and the E188 deficit (500 K vs 600–1000) is on the MC/entropy side. The quinary said the opposite (ratio 0.16). Two systems, two mechanisms: at MoNbTaVW the model's ordering scale is wrong; at Mo–Ta the scale is right and something in the sweep protocol (cell 6³, 100 samples, 24 T-steps, or the canonical ensemble's finite-size rounding) censors the peak. rand2 decides the branch; then E190b-on-v6b must be read per system, not as one number.

EXPERIMENTS.md · lines 12572–12582

E191 FINAL (2026-09-19 00:12) — N3 closed. Mo–Ta: E_B2 −184.5; randoms −123.1 / −92.4 / −95.7 (mean −103.7 ± 16.8) → ΔE_order(DFT) = +80.8 ± 16.8 vs v5 +80.3 — ratio 0.99. Branch: prediction 2, cleanly: v5's Mo–Ta ordering energy is right to the last meV, and its 500 K (E188) against the published 600–1000 K is not the model's ordering scale. The per-cell offsets are one-signed and large — v5 − DFT = +32.0 (B2), +31.8, +39.7, +23.2 (mean +31.7) — so v5 is uniformly ~32 meV too shallow at Mo₅₀Ta₅₀ in 16-atom cells, an absolute-energy error that cancels in ΔE_order and would not cancel in a formation-energy reward or a hull. Quinary: ΔE_order(DFT) +21.6 ± 1.8 vs v5 +3.6 on the same guessed cell (prediction 4 confirmed as a lower bound; the ratio is withdrawn, see the caveat above). Eight rows in data/dft/E191_*; they enter v6 with the site order fixed. What E200 now tests is the MC side for Mo–Ta; what E192 tests is the quinary on v5's own sampled states.

EXPERIMENTS.md · lines 13441–13445

E191's quinary reference was off by a factor of twenty. ΔE_order(v5) from E191's guessed ordered cell was +3.6; from the model's actual ground state it is +73.7. The "one sixth of DFT" reading was already withdrawn as a guess; this says how far off the guess was. E192's DFT verdict is unaffected — it was measured on MC-sampled cells, not on this one — and prediction (4) still has to run.

EXPERIMENTS.md · lines 13656–13719

E191's B2 Mo–Ta is −184.5 and E192's 54-atom Mo–Ta at the same 4×4×4 mesh is −184.49. Same crystal, same composition, same settings — so a 41 meV gap could not be cell size or k-mesh. It is the lattice constant: ground_state.py wrote cell.vasp straight from to_structure, which inherits pyeCE's PRIM constant a = 3.2935 Å. That is a mapping lattice for the cluster expansion, not a geometry. Every row in data/dft carries the convention "ideal bcc sites, Vegard a", and Vegard for Mo₅₀Ta₅₀ is 3.2460 Å. A 1.46 % expansion, and 41 meV/atom of it.

What this would have cost if it had gone unnoticed. The row was already written into data/dft/, and the v7 chain merges that whole directory into the training labels. A cell 41 meV off, labelled with a convention it does not have, would have gone into the first refit that was supposed to close the flywheel on the deployed model — and v7's own prediction (2), that the deep-row residual moves by less than 3 meV, would have been scored against a corrupted set.

Fixed, in this order: the row was moved out of data/dft to runs/quarantine_wrong_lattice/; the chain and the in-flight pw.x (MoNbTaVW, already running at 3.2935) were stopped; ground_state.py now rescales to BCC.cell_length(comp) before writing cell.vasp/pw.in and prints both constants; all three cells were regenerated (Mo–Ta 3.2460, MoNbTaVW 3.2049, MoNbTaW 3.2556) and re-verified as intact B2 and pseudo-binary B2; DFT relaunched on the corrected cells.

Prediction (4) is not scored yet. Nothing from the first DFT attempt counts. The −143.3 is retained only as the measurement of this defect.

Why it was caught. Not by a test, and not by the run itself — by asking why a number disagreed with a number for the same crystal already in the record. The same habit caught E213's scrambled site order an hour earlier. The v5 search results are unaffected: those are lattice-model evaluations keyed to the PRIM, and the ground states, spreads and ΔE_order values in E211 all stand.

And the defect is older than E211 — two rows already in the store have it (2026-09-21 08:3x). The new guard, run over every existing row, refuses E172_HfMoNbTaTiVWZr_s1.00 and E172b_HfMoNbTaTiVWZr_std: both sit at 3.2935 Å against a linear-Vegard 3.3426 for their composition, 1.47 % compressed, and both are stamped "ideal bcc sites, Vegard a".

But their energies are sound, and that is the more interesting finding. E172 ran a three-point volume scan on exactly that cell and recorded the minimum at a₀ = 3.2919 Å, with E(3.2935) − E(min) = +0.02 meV/atom. So the cell sits essentially on its own energy minimum; it is linear Vegard that is wrong here, by 1.5 %, on a composition spanning V (3.001 Å) to Zr (3.689 Å). Vegard's law is a proxy and large size mismatch breaks it.

So the rule is not "always Vegard", it is "no unstated mismatch". The guard now takes --allow-off-vegard <reason> and every stored row records a_conv_A, a_vegard_A and the reason, so no future row has to be trusted on its stamp. E172/E172b keep their energies and need only a corrected annotation; E211's Mo–Ta cell is a different case entirely — 3.2935 is not near its minimum (Mo–Ta's is at its Vegard 3.2460, confirmed by E191 matching RHEA's VASP row to 0.6 meV and Widom's −186), so its 41 meV really was an error.

The distinction that matters for the flywheel: a rung-4 cell must sit at its own energy minimum, because that is the geometry the eCE's labels correspond to. Linear Vegard is the cheap proxy for it, validated on Mo–Ta and falsified on the Hf/Zr-rich octonary. Any composition with that kind of size mismatch needs a volume scan before its DFT label is trusted — and there is no record of one for any cell other than E172's.

A second-order trap in the same repair (2026-09-21 08:4x). The relaunched DFT chain skipped Mo–Ta. Its guard is grep -q "JOB DONE" pw.out && continue, and the stale 3.2935 Å output still sat there carrying JOB DONE — the earlier attempt to move it had been part of a command that was refused, and I did not re-check before relaunching. So the one cell I most needed re-run was the one the chain declined to touch, and its old pw.out would have been read as the new result. Caught by noticing the "new" run reported a byte-identical total energy (−29458.98113510 Ry). The file is quarantined and runs/e211_mota_rerun.sh is queued behind the other two. A resume guard keyed on a file the repair was supposed to delete will silently preserve exactly what the repair was for.

Related entries

Built with PRISMWebsite and visualizations made using Claude