Does refitting the energy model on the larger new training set give a better model?
No. Every new version over-bound the search's alloys (13.6–19.6 meV/atom error, bar 8), so the earlier model stays.
In the log: the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)
mixedDate 2026-09-23 15:1x, as written in the logrung 4 · DFT4 predictions · 3 result paragraphsEXPERIMENTS.md lines 16090–16108, lines 16109–16128, lines 16129–16141, lines 16168–16184, lines 16220–16239
What E238 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E238.svg).
Pre-registration
(1)
A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds
E223's random cells at mean |error| ≤ 8 (v5 6.3).
failsA meets E237's three targets but over-binds the random …
(2)
A beats B on the eleven cells by ≥ 2
(E233 moved 4.7 at embedding 3) without losing more than 1 on held-48.
falsified10 Å pairs did not beat 6 Å on the eleven cells (11 …
(3)
C is within 1.5 of A
on every metric (123 rows).
falsifiedas worded (C differs from A by 2 …
(4)
D misses the random-cell bar again (> 10): the miss is the
embedding's, not the labels'.
confirmede5 misses the random-cell bar again (13 …
The pre-registration, as written
E238 — the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)
Labels (runs/rhea_labels_final_gated_sites.jsonl, 5,836 rows): E231's refined cubic pass-1
(true lattice minimum) + E234b's non-cubic pass-1 (ideal-shape reference, A100; 2,020 rows) →
pass 2 with the refined MEDIAN pure references and the 0.10 Å strain cut (5,878) → the site gate's
42 far-displaced rows dropped → every row carries its own ideal sites (merged emission fix: no row
is written in file order, none is skipped). Arms, 3 seeds each (MPS), v5's other settings:A embedding 6, pairs 10 Å — the candidate; B embedding 6, pairs 6 Å — range on the new
labels; C as A, trusted rows only (5,713); D embedding 5, pairs 10 Å — re-tests E237's e5
miss on the new labels. Scored on the eleven unbiased QE cells, the four Mo–Ta cells, B2 Mo–Ta,
the same held-48 file as E237, seed sd, and E223's three DFT random cells.
Predictions. (1) A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds
E223's random cells at mean |error| ≤ 8 (v5 6.3). (2) A beats B on the eleven cells by ≥ 2
(E233 moved 4.7 at embedding 3) without losing more than 1 on held-48. (3) C is within 1.5 of A
on every metric (123 rows). (4) D misses the random-cell bar again (> 10): the miss is the
embedding's, not the labels'. Decision: ship A if (1) holds, unless another arm beats it on
BOTH the eleven cells and the random cells by more than the seed sd; if A fails (1), ship the arm
that passes both; if none does, rung 0/1 stay on E237's e6 ensemble (old labels), stated as such.
runs/e238_chain.sh, two fits at a time.
E238 addendum (15:1x, before any fit finished) — rung 1's Hamiltonian. Measured, exact evaluator,
one thread, 432 sites: 10 Å pairs (embedding 3) 449 moves/s, embedding 6 at 6 Å 1,709. Rung 1
on a three-member 10 Å ensemble is ~1.8 M moves at ~150/s ≈ 3–4 h per composition; the 6 Å
ensemble ≈ 50 min. Rule: rung 0 uses the arm the decision above ships; rung 1 uses arm B (6 Å)
instead of A if B is within one seed sd of A on the ordering-sensitive rows — the four Mo–Ta cells'
mean, B2 Mo–Ta and the two E191 MoNbTaVW ordered cells — otherwise rung 1 uses A and the re-walk is
sized for it. Either way rung1_model names what was sampled.
Formal checks (Lean 4 + Mathlib, formal/, 15:xx–16:xx). The identities the ladder's numbers rest
on are now machine-checked, each tied to a pytest of the production function (formal/README.md):
the 2T temperature convention (pyeCE's acceptance at T ⇔ the true chain at s·T; transition_ece
and hysteresis return s × nominal), detailed balance and stationarity of sweep's one-attempt
kernel (checked numerically on the exact 4-site transition matrix to 1e-15), the SRO decode (exact
for α < 0; no decoder exists for α ≥ 0 — pyeCE's real encoder used), Warren–Cowley symmetry and
the bound α ≥ max(1 − 1/x_a, 1 − 1/x_b) with equality iff perfect order, the labeller's volume miss
(within the searched ±4 % bracket a miss only raises a label; an edge minimum is flagged
vertex_interior = false, marked untrusted, and — by default — still trained on), and the exact
evaluator's feature fold. Adversarial review of the STATEMENTS (fidelity, vacuity via compiled
witnesses, 4/4 code mutations caught) found two overclaims, both fixed. No proof found the code
disagreeing with the mathematics. All 5,836 E238 labels have an interior minimum.
Results
EXPERIMENTS.md · line 16129
E238 early read and a POST-HOC diagnostic (16:1x). Seed 0 of arms A and B, scored before the
other seeds: both new-label fits carry a negative signed bias on the eleven QE cells (A −11.3, B
−8.3; E237's e6 seed 0 on the old labels +3.4) and over-bind E223's random cells (mean |error| A
16.9, B 16.8, old e6 seed 0 9.3). One seed each — the decision still waits for three. Cause
located in the labels: on the 4,134 rows shared with v5 the new labels equal E231's exactly and sit
9.3 meV/atom below v5's, almost all composition-linear — the median pure references against
v5's minimum (per unit fraction: Ti −33.5, Cr −28.1, Hf −22.0, others within 6). The pure evidence
is thin: Ti has 3 pure frames, Hf's 12 corrected frames spread 34 meV; the least-corrected frame
sits at the median for Hf and Ti, at the minimum for Cr. Diagnostic, declared post hoc and NOT a
decision arm: a 2×2 at embedding 6 / 6 Å, one seed per new cell — {v5's rows, all rows} ×
{minimum, median references}; the other two cells exist (E237's e6 seed 0; arm B seed 0). It says
whether the over-binding is the reference choice or the 1,702 added ordered rows.
runs/e238diag_chain.sh.
EXPERIMENTS.md · line 16168
E238 diagnostic RESULT (17:2x) — the over-binding is the median references, not the new rows.
Embedding 6, 6 Å, seed 0 each; E223's three DFT random cells, mean |error| (meV/atom):
Eleven QE cells: 7.1 / 8.2 / 6.2 / 9.4 (same order: old-min, old-med, new-min, new-med); held-48
10.6 / 10.0 / 7.8 / 13.4. With minimum references the new rows cost nothing on the random cells
and give the best held-48 of the four; with median references both row sets over-bind. The
ordering-sensitive rows move within seed noise (four Mo–Ta cells' mean +4.3 / +1.8 / −4.5 / −6.0 (−4.5 corrected 2026-09-24 00:09 from −3.5: runs/e238diag_report.txt per-cell −22 +4 −0 −0);
per-cell seed sd ~8), consistent with a composition-linear shift cancelling at fixed composition.
Single seeds, and these are test cells: this LOCATES the problem, it does not choose the reference
— E239's anchor cells do that, independently. Consequence now: every E238 arm carries median
references, so E238's pre-registered fallback applies until E239 lands — rung 0/1 on E237's e6
ensemble (old labels), which is also the best measured model on the search's random cells (6.1).
EXPERIMENTS.md · line 16220
E238 RESULT (20:2x) — no new-label arm ships; the fallback stands. Three-seed ensembles
(runs/e238_report.txt), meV/atom:
(1) FAILS — A meets E237's three targets but over-binds the random cells by 19.6 (bar 8).
(2) FALSIFIED — 10 Å pairs did not beat 6 Å on the eleven cells (11.0 vs 6.5) on these labels;
they did on held-48 (8.6 vs 13.8). (3) FALSIFIED as worded (C differs from A by 2.6–3.7 on three
metrics; ensemble-mean noise is ~4 per cell, so this is not evidence that trust matters).
(4) CONFIRMED — e5 misses the random-cell bar again (13.6). Decision (pre-registered
fallback): no arm passes both faces, so rung 0/1 stay on E237's e6 ensemble on the old labels
— already the model E240 walks with. Cause, from the diagnostic: every arm carries the median
Ti/Cr/Hf references; E239 measures the right ones, after which the new labels (which cost nothing on
the random cells with minimum references and gave the best held-48) are refitted and re-tested.
The full record
This entry is written in 5 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 16090–16108
E238 — the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)
Labels (runs/rhea_labels_final_gated_sites.jsonl, 5,836 rows): E231's refined cubic pass-1
(true lattice minimum) + E234b's non-cubic pass-1 (ideal-shape reference, A100; 2,020 rows) →
pass 2 with the refined MEDIAN pure references and the 0.10 Å strain cut (5,878) → the site gate's
42 far-displaced rows dropped → every row carries its own ideal sites (merged emission fix: no row
is written in file order, none is skipped). Arms, 3 seeds each (MPS), v5's other settings:A embedding 6, pairs 10 Å — the candidate; B embedding 6, pairs 6 Å — range on the new
labels; C as A, trusted rows only (5,713); D embedding 5, pairs 10 Å — re-tests E237's e5
miss on the new labels. Scored on the eleven unbiased QE cells, the four Mo–Ta cells, B2 Mo–Ta,
the same held-48 file as E237, seed sd, and E223's three DFT random cells.
Predictions. (1) A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds
E223's random cells at mean |error| ≤ 8 (v5 6.3). (2) A beats B on the eleven cells by ≥ 2
(E233 moved 4.7 at embedding 3) without losing more than 1 on held-48. (3) C is within 1.5 of A
on every metric (123 rows). (4) D misses the random-cell bar again (> 10): the miss is the
embedding's, not the labels'. Decision: ship A if (1) holds, unless another arm beats it on
BOTH the eleven cells and the random cells by more than the seed sd; if A fails (1), ship the arm
that passes both; if none does, rung 0/1 stay on E237's e6 ensemble (old labels), stated as such.
runs/e238_chain.sh, two fits at a time.
EXPERIMENTS.md · lines 16109–16128
E238 addendum (15:1x, before any fit finished) — rung 1's Hamiltonian. Measured, exact evaluator,
one thread, 432 sites: 10 Å pairs (embedding 3) 449 moves/s, embedding 6 at 6 Å 1,709. Rung 1
on a three-member 10 Å ensemble is ~1.8 M moves at ~150/s ≈ 3–4 h per composition; the 6 Å
ensemble ≈ 50 min. Rule: rung 0 uses the arm the decision above ships; rung 1 uses arm B (6 Å)
instead of A if B is within one seed sd of A on the ordering-sensitive rows — the four Mo–Ta cells'
mean, B2 Mo–Ta and the two E191 MoNbTaVW ordered cells — otherwise rung 1 uses A and the re-walk is
sized for it. Either way rung1_model names what was sampled.
Formal checks (Lean 4 + Mathlib, formal/, 15:xx–16:xx). The identities the ladder's numbers rest
on are now machine-checked, each tied to a pytest of the production function (formal/README.md):
the 2T temperature convention (pyeCE's acceptance at T ⇔ the true chain at s·T; transition_ece
and hysteresis return s × nominal), detailed balance and stationarity of sweep's one-attempt
kernel (checked numerically on the exact 4-site transition matrix to 1e-15), the SRO decode (exact
for α < 0; no decoder exists for α ≥ 0 — pyeCE's real encoder used), Warren–Cowley symmetry and
the bound α ≥ max(1 − 1/x_a, 1 − 1/x_b) with equality iff perfect order, the labeller's volume miss
(within the searched ±4 % bracket a miss only raises a label; an edge minimum is flagged
vertex_interior = false, marked untrusted, and — by default — still trained on), and the exact
evaluator's feature fold. Adversarial review of the STATEMENTS (fidelity, vacuity via compiled
witnesses, 4/4 code mutations caught) found two overclaims, both fixed. No proof found the code
disagreeing with the mathematics. All 5,836 E238 labels have an interior minimum.
EXPERIMENTS.md · lines 16129–16141
E238 early read and a POST-HOC diagnostic (16:1x). Seed 0 of arms A and B, scored before the
other seeds: both new-label fits carry a negative signed bias on the eleven QE cells (A −11.3, B
−8.3; E237's e6 seed 0 on the old labels +3.4) and over-bind E223's random cells (mean |error| A
16.9, B 16.8, old e6 seed 0 9.3). One seed each — the decision still waits for three. Cause
located in the labels: on the 4,134 rows shared with v5 the new labels equal E231's exactly and sit
9.3 meV/atom below v5's, almost all composition-linear — the median pure references against
v5's minimum (per unit fraction: Ti −33.5, Cr −28.1, Hf −22.0, others within 6). The pure evidence
is thin: Ti has 3 pure frames, Hf's 12 corrected frames spread 34 meV; the least-corrected frame
sits at the median for Hf and Ti, at the minimum for Cr. Diagnostic, declared post hoc and NOT a
decision arm: a 2×2 at embedding 6 / 6 Å, one seed per new cell — {v5's rows, all rows} ×
{minimum, median references}; the other two cells exist (E237's e6 seed 0; arm B seed 0). It says
whether the over-binding is the reference choice or the 1,702 added ordered rows.
runs/e238diag_chain.sh.
EXPERIMENTS.md · lines 16168–16184
E238 diagnostic RESULT (17:2x) — the over-binding is the median references, not the new rows.
Embedding 6, 6 Å, seed 0 each; E223's three DFT random cells, mean |error| (meV/atom):
Eleven QE cells: 7.1 / 8.2 / 6.2 / 9.4 (same order: old-min, old-med, new-min, new-med); held-48
10.6 / 10.0 / 7.8 / 13.4. With minimum references the new rows cost nothing on the random cells
and give the best held-48 of the four; with median references both row sets over-bind. The
ordering-sensitive rows move within seed noise (four Mo–Ta cells' mean +4.3 / +1.8 / −4.5 / −6.0 (−4.5 corrected 2026-09-24 00:09 from −3.5: runs/e238diag_report.txt per-cell −22 +4 −0 −0);
per-cell seed sd ~8), consistent with a composition-linear shift cancelling at fixed composition.
Single seeds, and these are test cells: this LOCATES the problem, it does not choose the reference
— E239's anchor cells do that, independently. Consequence now: every E238 arm carries median
references, so E238's pre-registered fallback applies until E239 lands — rung 0/1 on E237's e6
ensemble (old labels), which is also the best measured model on the search's random cells (6.1).
EXPERIMENTS.md · lines 16220–16239
E238 RESULT (20:2x) — no new-label arm ships; the fallback stands. Three-seed ensembles
(runs/e238_report.txt), meV/atom:
(1) FAILS — A meets E237's three targets but over-binds the random cells by 19.6 (bar 8).
(2) FALSIFIED — 10 Å pairs did not beat 6 Å on the eleven cells (11.0 vs 6.5) on these labels;
they did on held-48 (8.6 vs 13.8). (3) FALSIFIED as worded (C differs from A by 2.6–3.7 on three
metrics; ensemble-mean noise is ~4 per cell, so this is not evidence that trust matters).
(4) CONFIRMED — e5 misses the random-cell bar again (13.6). Decision (pre-registered
fallback): no arm passes both faces, so rung 0/1 stay on E237's e6 ensemble on the old labels
— already the model E240 walks with. Cause, from the diagnostic: every arm carries the median
Ti/Cr/Hf references; E239 measures the right ones, after which the new labels (which cost nothing on
the random cells with minimum references and gave the best held-48) are refitted and re-tested.
Related entries
E231 — the E228 relabel, one variable: v5's rows, v5's recipe, corrected labels (pre-registered…
E234b — named in the log, no entry of its own
E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)
E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain…
E233 — the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)