Experiments · E238

Does refitting the energy model on the larger new training set give a better model?

No. Every new version over-bound the search's alloys (13.6–19.6 meV/atom error, bar 8), so the earlier model stays.

In the log: the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)

mixedDate 2026-09-23 15:1x, as written in the logrung 4 · DFT4 predictions · 3 result paragraphsEXPERIMENTS.md lines 16090–16108, lines 16109–16128, lines 16129–16141, lines 16168–16184, lines 16220–16239
exp E238 diagram
What E238 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E238.svg).

Pre-registration

  1. (1)
    A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds E223's random cells at mean |error| ≤ 8 (v5 6.3).
    failsA meets E237's three targets but over-binds the random …
  2. (2)
    A beats B on the eleven cells by ≥ 2 (E233 moved 4.7 at embedding 3) without losing more than 1 on held-48.
    falsified10 Å pairs did not beat 6 Å on the eleven cells (11 …
  3. (3)
    C is within 1.5 of A on every metric (123 rows).
    falsifiedas worded (C differs from A by 2 …
  4. (4)
    D misses the random-cell bar again (> 10): the miss is the embedding's, not the labels'.
    confirmede5 misses the random-cell bar again (13 …
The pre-registration, as written

E238 — the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)

Labels (runs/rhea_labels_final_gated_sites.jsonl, 5,836 rows): E231's refined cubic pass-1 (true lattice minimum) + E234b's non-cubic pass-1 (ideal-shape reference, A100; 2,020 rows) → pass 2 with the refined MEDIAN pure references and the 0.10 Å strain cut (5,878) → the site gate's 42 far-displaced rows dropped → every row carries its own ideal sites (merged emission fix: no row is written in file order, none is skipped). Arms, 3 seeds each (MPS), v5's other settings: A embedding 6, pairs 10 Å — the candidate; B embedding 6, pairs 6 Å — range on the new labels; C as A, trusted rows only (5,713); D embedding 5, pairs 10 Å — re-tests E237's e5 miss on the new labels. Scored on the eleven unbiased QE cells, the four Mo–Ta cells, B2 Mo–Ta, the same held-48 file as E237, seed sd, and E223's three DFT random cells. Predictions. (1) A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds E223's random cells at mean |error| ≤ 8 (v5 6.3). (2) A beats B on the eleven cells by ≥ 2 (E233 moved 4.7 at embedding 3) without losing more than 1 on held-48. (3) C is within 1.5 of A on every metric (123 rows). (4) D misses the random-cell bar again (> 10): the miss is the embedding's, not the labels'. Decision: ship A if (1) holds, unless another arm beats it on BOTH the eleven cells and the random cells by more than the seed sd; if A fails (1), ship the arm that passes both; if none does, rung 0/1 stay on E237's e6 ensemble (old labels), stated as such. runs/e238_chain.sh, two fits at a time.

E238 addendum (15:1x, before any fit finished) — rung 1's Hamiltonian. Measured, exact evaluator, one thread, 432 sites: 10 Å pairs (embedding 3) 449 moves/s, embedding 6 at 6 Å 1,709. Rung 1 on a three-member 10 Å ensemble is ~1.8 M moves at ~150/s ≈ 3–4 h per composition; the 6 Å ensemble ≈ 50 min. Rule: rung 0 uses the arm the decision above ships; rung 1 uses arm B (6 Å) instead of A if B is within one seed sd of A on the ordering-sensitive rows — the four Mo–Ta cells' mean, B2 Mo–Ta and the two E191 MoNbTaVW ordered cells — otherwise rung 1 uses A and the re-walk is sized for it. Either way rung1_model names what was sampled.

Formal checks (Lean 4 + Mathlib, formal/, 15:xx–16:xx). The identities the ladder's numbers rest on are now machine-checked, each tied to a pytest of the production function (formal/README.md): the 2T temperature convention (pyeCE's acceptance at T ⇔ the true chain at s·T; transition_ece and hysteresis return s × nominal), detailed balance and stationarity of sweep's one-attempt kernel (checked numerically on the exact 4-site transition matrix to 1e-15), the SRO decode (exact for α < 0; no decoder exists for α ≥ 0 — pyeCE's real encoder used), Warren–Cowley symmetry and the bound α ≥ max(1 − 1/x_a, 1 − 1/x_b) with equality iff perfect order, the labeller's volume miss (within the searched ±4 % bracket a miss only raises a label; an edge minimum is flagged vertex_interior = false, marked untrusted, and — by default — still trained on), and the exact evaluator's feature fold. Adversarial review of the STATEMENTS (fidelity, vacuity via compiled witnesses, 4/4 code mutations caught) found two overclaims, both fixed. No proof found the code disagreeing with the mathematics. All 5,836 E238 labels have an interior minimum.

Results

EXPERIMENTS.md · line 16129

E238 early read and a POST-HOC diagnostic (16:1x). Seed 0 of arms A and B, scored before the other seeds: both new-label fits carry a negative signed bias on the eleven QE cells (A −11.3, B −8.3; E237's e6 seed 0 on the old labels +3.4) and over-bind E223's random cells (mean |error| A 16.9, B 16.8, old e6 seed 0 9.3). One seed each — the decision still waits for three. Cause located in the labels: on the 4,134 rows shared with v5 the new labels equal E231's exactly and sit 9.3 meV/atom below v5's, almost all composition-linear — the median pure references against v5's minimum (per unit fraction: Ti −33.5, Cr −28.1, Hf −22.0, others within 6). The pure evidence is thin: Ti has 3 pure frames, Hf's 12 corrected frames spread 34 meV; the least-corrected frame sits at the median for Hf and Ti, at the minimum for Cr. Diagnostic, declared post hoc and NOT a decision arm: a 2×2 at embedding 6 / 6 Å, one seed per new cell — {v5's rows, all rows} × {minimum, median references}; the other two cells exist (E237's e6 seed 0; arm B seed 0). It says whether the over-binding is the reference choice or the 1,702 added ordered rows. runs/e238diag_chain.sh.

EXPERIMENTS.md · line 16168

E238 diagnostic RESULT (17:2x) — the over-binding is the median references, not the new rows. Embedding 6, 6 Å, seed 0 each; E223's three DFT random cells, mean |error| (meV/atom):

minimum references median references
v5's rows 9.3 (E237 e6 s0) 23.9
all rows (+1,702 ordered) 9.7 16.8 (arm B s0)

Eleven QE cells: 7.1 / 8.2 / 6.2 / 9.4 (same order: old-min, old-med, new-min, new-med); held-48 10.6 / 10.0 / 7.8 / 13.4. With minimum references the new rows cost nothing on the random cells and give the best held-48 of the four; with median references both row sets over-bind. The ordering-sensitive rows move within seed noise (four Mo–Ta cells' mean +4.3 / +1.8 / −4.5 / −6.0 (−4.5 corrected 2026-09-24 00:09 from −3.5: runs/e238diag_report.txt per-cell −22 +4 −0 −0); per-cell seed sd ~8), consistent with a composition-linear shift cancelling at fixed composition. Single seeds, and these are test cells: this LOCATES the problem, it does not choose the reference — E239's anchor cells do that, independently. Consequence now: every E238 arm carries median references, so E238's pre-registered fallback applies until E239 lands — rung 0/1 on E237's e6 ensemble (old labels), which is also the best measured model on the search's random cells (6.1).

EXPERIMENTS.md · line 16220

E238 RESULT (20:2x) — no new-label arm ships; the fallback stands. Three-seed ensembles (runs/e238_report.txt), meV/atom:

arm eleven QE cells four Mo–Ta mean B2 Mo–Ta held-48 E223 random cells, mean abs.
A e6, 10 Å 11.0 −7.0 −6.0 8.6 19.6
B e6, 6 Å 6.5 −2.5 +2.9 13.8 19.3
C e6, 10 Å, trusted only 8.2 −2.5 −2.3 11.2 18.5
D e5, 10 Å 11.2 −7.0 −14.4 11.7 13.6
E237 e6 (old labels) 7.5 +5.5 −5.2 7.9 6.1

(1) FAILS — A meets E237's three targets but over-binds the random cells by 19.6 (bar 8). (2) FALSIFIED — 10 Å pairs did not beat 6 Å on the eleven cells (11.0 vs 6.5) on these labels; they did on held-48 (8.6 vs 13.8). (3) FALSIFIED as worded (C differs from A by 2.6–3.7 on three metrics; ensemble-mean noise is ~4 per cell, so this is not evidence that trust matters). (4) CONFIRMED — e5 misses the random-cell bar again (13.6). Decision (pre-registered fallback): no arm passes both faces, so rung 0/1 stay on E237's e6 ensemble on the old labels — already the model E240 walks with. Cause, from the diagnostic: every arm carries the median Ti/Cr/Hf references; E239 measures the right ones, after which the new labels (which cost nothing on the random cells with minimum references and gave the best held-48) are refitted and re-tested.

The full record

This entry is written in 5 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 16090–16108

E238 — the rung-0/1 model that ships (pre-registered 2026-09-23 15:1x)

Labels (runs/rhea_labels_final_gated_sites.jsonl, 5,836 rows): E231's refined cubic pass-1 (true lattice minimum) + E234b's non-cubic pass-1 (ideal-shape reference, A100; 2,020 rows) → pass 2 with the refined MEDIAN pure references and the 0.10 Å strain cut (5,878) → the site gate's 42 far-displaced rows dropped → every row carries its own ideal sites (merged emission fix: no row is written in file order, none is skipped). Arms, 3 seeds each (MPS), v5's other settings: A embedding 6, pairs 10 Å — the candidate; B embedding 6, pairs 6 Å — range on the new labels; C as A, trusted rows only (5,713); D embedding 5, pairs 10 Å — re-tests E237's e5 miss on the new labels. Scored on the eleven unbiased QE cells, the four Mo–Ta cells, B2 Mo–Ta, the same held-48 file as E237, seed sd, and E223's three DFT random cells. Predictions. (1) A meets E237's targets (MAE ≤ 14, Mo–Ta mean < +15, held-48 ≤ 9) AND holds E223's random cells at mean |error| ≤ 8 (v5 6.3). (2) A beats B on the eleven cells by ≥ 2 (E233 moved 4.7 at embedding 3) without losing more than 1 on held-48. (3) C is within 1.5 of A on every metric (123 rows). (4) D misses the random-cell bar again (> 10): the miss is the embedding's, not the labels'. Decision: ship A if (1) holds, unless another arm beats it on BOTH the eleven cells and the random cells by more than the seed sd; if A fails (1), ship the arm that passes both; if none does, rung 0/1 stay on E237's e6 ensemble (old labels), stated as such. runs/e238_chain.sh, two fits at a time.

EXPERIMENTS.md · lines 16109–16128

E238 addendum (15:1x, before any fit finished) — rung 1's Hamiltonian. Measured, exact evaluator, one thread, 432 sites: 10 Å pairs (embedding 3) 449 moves/s, embedding 6 at 6 Å 1,709. Rung 1 on a three-member 10 Å ensemble is ~1.8 M moves at ~150/s ≈ 3–4 h per composition; the 6 Å ensemble ≈ 50 min. Rule: rung 0 uses the arm the decision above ships; rung 1 uses arm B (6 Å) instead of A if B is within one seed sd of A on the ordering-sensitive rows — the four Mo–Ta cells' mean, B2 Mo–Ta and the two E191 MoNbTaVW ordered cells — otherwise rung 1 uses A and the re-walk is sized for it. Either way rung1_model names what was sampled.

Formal checks (Lean 4 + Mathlib, formal/, 15:xx–16:xx). The identities the ladder's numbers rest on are now machine-checked, each tied to a pytest of the production function (formal/README.md): the 2T temperature convention (pyeCE's acceptance at T ⇔ the true chain at s·T; transition_ece and hysteresis return s × nominal), detailed balance and stationarity of sweep's one-attempt kernel (checked numerically on the exact 4-site transition matrix to 1e-15), the SRO decode (exact for α < 0; no decoder exists for α ≥ 0 — pyeCE's real encoder used), Warren–Cowley symmetry and the bound α ≥ max(1 − 1/x_a, 1 − 1/x_b) with equality iff perfect order, the labeller's volume miss (within the searched ±4 % bracket a miss only raises a label; an edge minimum is flagged vertex_interior = false, marked untrusted, and — by default — still trained on), and the exact evaluator's feature fold. Adversarial review of the STATEMENTS (fidelity, vacuity via compiled witnesses, 4/4 code mutations caught) found two overclaims, both fixed. No proof found the code disagreeing with the mathematics. All 5,836 E238 labels have an interior minimum.

EXPERIMENTS.md · lines 16129–16141

E238 early read and a POST-HOC diagnostic (16:1x). Seed 0 of arms A and B, scored before the other seeds: both new-label fits carry a negative signed bias on the eleven QE cells (A −11.3, B −8.3; E237's e6 seed 0 on the old labels +3.4) and over-bind E223's random cells (mean |error| A 16.9, B 16.8, old e6 seed 0 9.3). One seed each — the decision still waits for three. Cause located in the labels: on the 4,134 rows shared with v5 the new labels equal E231's exactly and sit 9.3 meV/atom below v5's, almost all composition-linear — the median pure references against v5's minimum (per unit fraction: Ti −33.5, Cr −28.1, Hf −22.0, others within 6). The pure evidence is thin: Ti has 3 pure frames, Hf's 12 corrected frames spread 34 meV; the least-corrected frame sits at the median for Hf and Ti, at the minimum for Cr. Diagnostic, declared post hoc and NOT a decision arm: a 2×2 at embedding 6 / 6 Å, one seed per new cell — {v5's rows, all rows} × {minimum, median references}; the other two cells exist (E237's e6 seed 0; arm B seed 0). It says whether the over-binding is the reference choice or the 1,702 added ordered rows. runs/e238diag_chain.sh.

EXPERIMENTS.md · lines 16168–16184

E238 diagnostic RESULT (17:2x) — the over-binding is the median references, not the new rows. Embedding 6, 6 Å, seed 0 each; E223's three DFT random cells, mean |error| (meV/atom):

minimum references median references
v5's rows 9.3 (E237 e6 s0) 23.9
all rows (+1,702 ordered) 9.7 16.8 (arm B s0)

Eleven QE cells: 7.1 / 8.2 / 6.2 / 9.4 (same order: old-min, old-med, new-min, new-med); held-48 10.6 / 10.0 / 7.8 / 13.4. With minimum references the new rows cost nothing on the random cells and give the best held-48 of the four; with median references both row sets over-bind. The ordering-sensitive rows move within seed noise (four Mo–Ta cells' mean +4.3 / +1.8 / −4.5 / −6.0 (−4.5 corrected 2026-09-24 00:09 from −3.5: runs/e238diag_report.txt per-cell −22 +4 −0 −0); per-cell seed sd ~8), consistent with a composition-linear shift cancelling at fixed composition. Single seeds, and these are test cells: this LOCATES the problem, it does not choose the reference — E239's anchor cells do that, independently. Consequence now: every E238 arm carries median references, so E238's pre-registered fallback applies until E239 lands — rung 0/1 on E237's e6 ensemble (old labels), which is also the best measured model on the search's random cells (6.1).

EXPERIMENTS.md · lines 16220–16239

E238 RESULT (20:2x) — no new-label arm ships; the fallback stands. Three-seed ensembles (runs/e238_report.txt), meV/atom:

arm eleven QE cells four Mo–Ta mean B2 Mo–Ta held-48 E223 random cells, mean abs.
A e6, 10 Å 11.0 −7.0 −6.0 8.6 19.6
B e6, 6 Å 6.5 −2.5 +2.9 13.8 19.3
C e6, 10 Å, trusted only 8.2 −2.5 −2.3 11.2 18.5
D e5, 10 Å 11.2 −7.0 −14.4 11.7 13.6
E237 e6 (old labels) 7.5 +5.5 −5.2 7.9 6.1

(1) FAILS — A meets E237's three targets but over-binds the random cells by 19.6 (bar 8). (2) FALSIFIED — 10 Å pairs did not beat 6 Å on the eleven cells (11.0 vs 6.5) on these labels; they did on held-48 (8.6 vs 13.8). (3) FALSIFIED as worded (C differs from A by 2.6–3.7 on three metrics; ensemble-mean noise is ~4 per cell, so this is not evidence that trust matters). (4) CONFIRMED — e5 misses the random-cell bar again (13.6). Decision (pre-registered fallback): no arm passes both faces, so rung 0/1 stay on E237's e6 ensemble on the old labels — already the model E240 walks with. Cause, from the diagnostic: every arm carries the median Ti/Cr/Hf references; E239 measures the right ones, after which the new labels (which cost nothing on the random cells with minimum references and gave the best held-48) are refitted and re-tested.

Related entries

Built with PRISMWebsite and visualizations made using Claude