Experiments · E223

Can alloys proposed by a search pass the corrected ladder's ordering and stability checks?

No. Of 32 finds, 24 could not pass because the simulation stopped at 100 K; none reached the 0.8 success bar.

In the log: the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)

supersededDate 2026-09-22 04:2x, as written in the logrung 4 · DFT3 predictions · 1 result paragraphEXPERIMENTS.md lines 15451–15489, lines 16007–16036, lines 16353–16360
exp E223 diagram
What E223 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E223.svg).

Pre-registration

  1. (1)
    Of 32 leaders, 16 ± 6 pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing high falsifies the reading that the ×2 correction moves search finds up as it moved the scorecard.
    no verdict written against it
  2. (2)
    The fly's leaders do not differ from elitist's in element support (mean element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4 of 32).
    no verdict written against it
  3. (3)
    Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤ −40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk ~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s; random_cell.py on a pick reproduced E222's own numbers.
    no verdict written against it
The pre-registration, as written

E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)

Gap 2: rung 1 became the Metropolis sampler on the physical axis on 2026-09-21 and has since run only on hand-picked systems (E210b, E190b's re-read, E222's cell). No composition a SEARCH proposed has ever been through it. runs/e223_chain.sh (referee-written, verified by running its parts): one free search on E201b's arms and budget (elitist, fly; 200 × 4 × 2 seeds; pure reward; the fb,oa,5ht brain flags) with FORAGER_RUNG0=ece and no --confirm (rung 1 inline is 20–32 min per qualifier), then scripts/search/walk_finds.py (new) walks the pooled niche leaders — de-duplicated on composition, capped at 32 — up rung 1 and rung 2 (MACE hull), one core, ~25 min each, each result on disk before the next. Rung 3 is not walked while trust_kinetics is False. The top 3 by ladder p not already in data/dft are written as 54-atom cells to runs/e223_dft_queue.txt; pw.x is not launched. Design: docs/research/2026-09-22-e223-first-search-on-corrected-ladder.md. One deviation from E201b: ce_12element restricted by FORAGER_ELEMENTS to the eCE's nine, so every leader is walkable and Cr — where v5 and SOB20 disagree — is in the space.

Two ladder defects found while wiring it, both fixed before the chain was armed. (i) ECEVerifier was not callable: Verifier.__call__ implements the walk, the wrapper delegates by __getattr__, and Python looks special methods up on the type — V(x) raised, and V.base(x) would silently have walked from the MACE-frame screen. ECEVerifier.__call__ now runs the base walk with the wrapper's own screen. (ii) A censored rung 1 reached rung 2 as T = NaN and came out as p = NaN — neither pass nor fail; Verifier._requirement now returns the safe zero with rung1_status at rungs 2 and 3. Both pinned in tests/test_rung1_metropolis.py. Consequence the referee measured: with a 227 K grid floor the ladder can pass a composition only on the HIGH side of the window; "below 90 K" is unreachable at this budget.

Predictions (on the ladder's own convention — cooling leg, largest peak: E210b reads 747 / 463 / 1163 / 1210 K that way, two of four above 1000). (1) Of 32 leaders, 16 ± 6 pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing high falsifies the reading that the ×2 correction moves search finds up as it moved the scorecard. (2) The fly's leaders do not differ from elitist's in element support (mean element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4 of 32). (3) Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤ −40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk ~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s; random_cell.py on a pick reproduced E222's own numbers.

E223 scored — the walk it pre-registered. runs/e223_walk/summary.jsonl (32 leaders; v5, old rung-1 floor). (1) Falsified: 1 of 32 passes high (Mo₇₀Hf₂₀Ti₁₀, 1291 K) against 16 ± 6 (fewer than 6 falsifies); 7 order inside the window (424–896 K), 24 show no transition above 100 K. (2) Falsified: the arms differ in support: fly 3.17 elements per leader, elitist 5.21 (elements ≥ 1 at.%), and 1 of 21 supports is reached by both (≥ 60 % predicted); Cr is in 4 of 32 (at most 4: holds). (3) Neither branch: best finite p 0.572 (Mo₄₂Ta₂₉Ti₂₉, drive_cold −114), below the 0.8 success bar and above the 0.5 falsifier. Superseded as a verdict on the finds by E236/E240, which re-walked the same 32 on the corrected ladder (e6, floor 60 K).

Results

EXPERIMENTS.md · line 16007

E223 rung-4 RESULT (14:1x, Colab A100) — the picks' random cells. DFT formation energy of each pick's random 54-atom cell (v3 standard) against v5's recorded prediction and the two ensembles on the identical cell: Mo₄₂Ta₂₉Ti₂₉ −56.5 (v5 −56.9; v5×3 −50.9 ± 7.7; e9×3 −69.5 ± 1.2); Mo₇₀Hf₂₀Ti₁₀ −31.2 (−29.6; −24.9 ± 6.0; −33.5 ± 8.5); Mo₄₁Ti₂₅Ta₁₀Nb₈W₇Hf₄Zr₄ −14.3 (−29.8; −21.3 ± 8.0; −27.2 ± 1.8). Mean |error|: v5 5.8, v5×3 6.3, e9×3 9.4 — and confidently wrong: −13 on two of three with a seed spread of 1.2–1.8. E229's two faces, now on the search's own targets: embedding 9 fixes ordered states and over-binds random solutions of unseen compositions, and its seed spread does not flag it. Rung 0 (random-solution energy) and rung 1 (ordering) now favour different models; the model that ships must pass both — E237's middle embeddings and E233's 10 Å pairs are the candidates. The random-cell energies of the three picks are confirmed at rung 4 by v5 to within 2 meV on two of three; their ORDERING remains the open question.

Rung-1 speed SOLVED (14:0x) — the exact evaluator. forager/ladder/ece_distill.py::ExactECE (Fable-written) re-evaluates pyeCE's forward exactly in numpy: per-orbit correlation tensors, the linear map folded into the first network layer, no per-cluster-function gather. Acceptance on the teacher runs/ece_v5e4_seed0, Mo₄₂Ta₂₉Ti₂₉, the production cell (432 sites) and settings (scripts/validate/distill_check.py, runs/distill_check/): per-swap energy differences against pyeCE's own forward MAE 0.0012 meV, max 0.0071 meV over 2,304 swaps whose mean size is 160–350 meV (float32 rounding), on configurations sampled by either Hamiltonian at 300–1500 K; rung 1 on the same seed 1148 ± 114 K against the teacher's 1175 ± 114 K (the difference is chain divergence from rounding, inside one grid step's σ). Throughput, swap_delta moves/s, one thread, 432 sites: v5 2,788 (pyeCE forward 1,009), embedding 4 1,987 (403), embedding 9 1,519 (28 — 54×). Rung-1 cost no longer depends materially on embedding size, so E237's decision reduces to accuracy: prediction (3) is moot, and the embedding is chosen by the DFT targets alone. The fitted alternative (DistilledCE, pairs + triplets distilled from the teacher) FAILS parity (fresh-state swap RMS 55.7 meV) although its T_c landed at 1170 K — a surrogate can hit one number by accident, which is why parity, not T_c, is the acceptance test; it is kept only as the recorded negative. Rung 1 in ECEVerifier now samples the MEAN of every ensemble member through the exact evaluator (_hamiltonian), so rung 0 and rung 1 score the same model; rung1_model names every member.

The full record

This entry is written in 3 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 15451–15489

E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)

Gap 2: rung 1 became the Metropolis sampler on the physical axis on 2026-09-21 and has since run only on hand-picked systems (E210b, E190b's re-read, E222's cell). No composition a SEARCH proposed has ever been through it. runs/e223_chain.sh (referee-written, verified by running its parts): one free search on E201b's arms and budget (elitist, fly; 200 × 4 × 2 seeds; pure reward; the fb,oa,5ht brain flags) with FORAGER_RUNG0=ece and no --confirm (rung 1 inline is 20–32 min per qualifier), then scripts/search/walk_finds.py (new) walks the pooled niche leaders — de-duplicated on composition, capped at 32 — up rung 1 and rung 2 (MACE hull), one core, ~25 min each, each result on disk before the next. Rung 3 is not walked while trust_kinetics is False. The top 3 by ladder p not already in data/dft are written as 54-atom cells to runs/e223_dft_queue.txt; pw.x is not launched. Design: docs/research/2026-09-22-e223-first-search-on-corrected-ladder.md. One deviation from E201b: ce_12element restricted by FORAGER_ELEMENTS to the eCE's nine, so every leader is walkable and Cr — where v5 and SOB20 disagree — is in the space.

Two ladder defects found while wiring it, both fixed before the chain was armed. (i) ECEVerifier was not callable: Verifier.__call__ implements the walk, the wrapper delegates by __getattr__, and Python looks special methods up on the type — V(x) raised, and V.base(x) would silently have walked from the MACE-frame screen. ECEVerifier.__call__ now runs the base walk with the wrapper's own screen. (ii) A censored rung 1 reached rung 2 as T = NaN and came out as p = NaN — neither pass nor fail; Verifier._requirement now returns the safe zero with rung1_status at rungs 2 and 3. Both pinned in tests/test_rung1_metropolis.py. Consequence the referee measured: with a 227 K grid floor the ladder can pass a composition only on the HIGH side of the window; "below 90 K" is unreachable at this budget.

Predictions (on the ladder's own convention — cooling leg, largest peak: E210b reads 747 / 463 / 1163 / 1210 K that way, two of four above 1000). (1) Of 32 leaders, 16 ± 6 pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing high falsifies the reading that the ×2 correction moves search finds up as it moved the scorecard. (2) The fly's leaders do not differ from elitist's in element support (mean element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4 of 32). (3) Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤ −40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk ~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s; random_cell.py on a pick reproduced E222's own numbers.

EXPERIMENTS.md · lines 16007–16036

E223 rung-4 RESULT (14:1x, Colab A100) — the picks' random cells. DFT formation energy of each pick's random 54-atom cell (v3 standard) against v5's recorded prediction and the two ensembles on the identical cell: Mo₄₂Ta₂₉Ti₂₉ −56.5 (v5 −56.9; v5×3 −50.9 ± 7.7; e9×3 −69.5 ± 1.2); Mo₇₀Hf₂₀Ti₁₀ −31.2 (−29.6; −24.9 ± 6.0; −33.5 ± 8.5); Mo₄₁Ti₂₅Ta₁₀Nb₈W₇Hf₄Zr₄ −14.3 (−29.8; −21.3 ± 8.0; −27.2 ± 1.8). Mean |error|: v5 5.8, v5×3 6.3, e9×3 9.4 — and confidently wrong: −13 on two of three with a seed spread of 1.2–1.8. E229's two faces, now on the search's own targets: embedding 9 fixes ordered states and over-binds random solutions of unseen compositions, and its seed spread does not flag it. Rung 0 (random-solution energy) and rung 1 (ordering) now favour different models; the model that ships must pass both — E237's middle embeddings and E233's 10 Å pairs are the candidates. The random-cell energies of the three picks are confirmed at rung 4 by v5 to within 2 meV on two of three; their ORDERING remains the open question.

Rung-1 speed SOLVED (14:0x) — the exact evaluator. forager/ladder/ece_distill.py::ExactECE (Fable-written) re-evaluates pyeCE's forward exactly in numpy: per-orbit correlation tensors, the linear map folded into the first network layer, no per-cluster-function gather. Acceptance on the teacher runs/ece_v5e4_seed0, Mo₄₂Ta₂₉Ti₂₉, the production cell (432 sites) and settings (scripts/validate/distill_check.py, runs/distill_check/): per-swap energy differences against pyeCE's own forward MAE 0.0012 meV, max 0.0071 meV over 2,304 swaps whose mean size is 160–350 meV (float32 rounding), on configurations sampled by either Hamiltonian at 300–1500 K; rung 1 on the same seed 1148 ± 114 K against the teacher's 1175 ± 114 K (the difference is chain divergence from rounding, inside one grid step's σ). Throughput, swap_delta moves/s, one thread, 432 sites: v5 2,788 (pyeCE forward 1,009), embedding 4 1,987 (403), embedding 9 1,519 (28 — 54×). Rung-1 cost no longer depends materially on embedding size, so E237's decision reduces to accuracy: prediction (3) is moot, and the embedding is chosen by the DFT targets alone. The fitted alternative (DistilledCE, pairs + triplets distilled from the teacher) FAILS parity (fresh-state swap RMS 55.7 meV) although its T_c landed at 1170 K — a surrogate can hit one number by accident, which is why parity, not T_c, is the acceptance test; it is kept only as the recorded negative. Rung 1 in ECEVerifier now samples the MEAN of every ensemble member through the exact evaluator (_hamiltonian), so rung 0 and rung 1 score the same model; rung1_model names every member.

EXPERIMENTS.md · lines 16353–16360

E223 scored — the walk it pre-registered. runs/e223_walk/summary.jsonl (32 leaders; v5, old rung-1 floor). (1) Falsified: 1 of 32 passes high (Mo₇₀Hf₂₀Ti₁₀, 1291 K) against 16 ± 6 (fewer than 6 falsifies); 7 order inside the window (424–896 K), 24 show no transition above 100 K. (2) Falsified: the arms differ in support: fly 3.17 elements per leader, elitist 5.21 (elements ≥ 1 at.%), and 1 of 21 supports is reached by both (≥ 60 % predicted); Cr is in 4 of 32 (at most 4: holds). (3) Neither branch: best finite p 0.572 (Mo₄₂Ta₂₉Ti₂₉, drive_cold −114), below the 0.8 success bar and above the 0.5 falsifier. Superseded as a verdict on the finds by E236/E240, which re-walked the same 32 on the corrected ladder (e6, floor 60 K).

Related entries

Built with PRISMWebsite and visualizations made using Claude