Can alloys proposed by a search pass the corrected ladder's ordering and stability checks?
No. Of 32 finds, 24 could not pass because the simulation stopped at 100 K; none reached the 0.8 success bar.
In the log: the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)
supersededDate 2026-09-22 04:2x, as written in the logrung 4 · DFT3 predictions · 1 result paragraphEXPERIMENTS.md lines 15451–15489, lines 16007–16036, lines 16353–16360
What E223 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E223.svg).
Pre-registration
(1)
Of 32 leaders, 16 ± 6
pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing
high falsifies the reading that the ×2 correction moves search finds up as it moved the
scorecard.
no verdict written against it
(2)
The fly's leaders do not differ from elitist's in element support (mean
element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4
of 32).
no verdict written against it
(3)
Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤
−40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored
NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk
~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on
Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s;
random_cell.py on a pick reproduced E222's own numbers.
no verdict written against it
The pre-registration, as written
E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)
Gap 2: rung 1 became the Metropolis sampler on the physical axis on 2026-09-21 and has since
run only on hand-picked systems (E210b, E190b's re-read, E222's cell). No composition a
SEARCH proposed has ever been through it.runs/e223_chain.sh (referee-written, verified
by running its parts): one free search on E201b's arms and budget (elitist, fly; 200 × 4 × 2
seeds; pure reward; the fb,oa,5ht brain flags) with FORAGER_RUNG0=ece and no --confirm
(rung 1 inline is 20–32 min per qualifier), then scripts/search/walk_finds.py (new) walks
the pooled niche leaders — de-duplicated on composition, capped at 32 — up rung 1 and rung 2
(MACE hull), one core, ~25 min each, each result on disk before the next. Rung 3 is not
walked while trust_kinetics is False. The top 3 by ladder p not already in data/dft are
written as 54-atom cells to runs/e223_dft_queue.txt; pw.x is not launched. Design:
docs/research/2026-09-22-e223-first-search-on-corrected-ladder.md. One deviation from
E201b: ce_12element restricted by FORAGER_ELEMENTS to the eCE's nine, so every leader is
walkable and Cr — where v5 and SOB20 disagree — is in the space.
Two ladder defects found while wiring it, both fixed before the chain was armed. (i)
ECEVerifier was not callable: Verifier.__call__ implements the walk, the wrapper delegates
by __getattr__, and Python looks special methods up on the type — V(x) raised, and
V.base(x) would silently have walked from the MACE-frame screen. ECEVerifier.__call__ now
runs the base walk with the wrapper's own screen. (ii) A censored rung 1 reached rung 2 as
T = NaN and came out as p = NaN — neither pass nor fail; Verifier._requirement now
returns the safe zero with rung1_status at rungs 2 and 3. Both pinned in
tests/test_rung1_metropolis.py. Consequence the referee measured: with a 227 K grid floor
the ladder can pass a composition only on the HIGH side of the window; "below 90 K" is
unreachable at this budget.
Predictions (on the ladder's own convention — cooling leg, largest peak: E210b reads
747 / 463 / 1163 / 1210 K that way, two of four above 1000). (1) Of 32 leaders, 16 ± 6
pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing
high falsifies the reading that the ×2 correction moves search finds up as it moved the
scorecard. (2) The fly's leaders do not differ from elitist's in element support (mean
element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4
of 32). (3) Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤
−40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored
NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk
~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on
Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s;
random_cell.py on a pick reproduced E222's own numbers.
E223 scored — the walk it pre-registered.runs/e223_walk/summary.jsonl (32 leaders; v5,
old rung-1 floor). (1) Falsified: 1 of 32 passes high (Mo₇₀Hf₂₀Ti₁₀, 1291 K) against 16 ± 6
(fewer than 6 falsifies); 7 order inside the window (424–896 K), 24 show no transition above
100 K. (2) Falsified: the arms differ in support: fly 3.17 elements per leader, elitist 5.21
(elements ≥ 1 at.%), and 1 of 21 supports is reached by both (≥ 60 % predicted); Cr is in 4 of
32 (at most 4: holds). (3) Neither branch: best finite p 0.572 (Mo₄₂Ta₂₉Ti₂₉, drive_cold
−114), below the 0.8 success bar and above the 0.5 falsifier. Superseded as a verdict on the
finds by E236/E240, which re-walked the same 32 on the corrected ladder (e6, floor 60 K).
Results
EXPERIMENTS.md · line 16007
E223 rung-4 RESULT (14:1x, Colab A100) — the picks' random cells. DFT formation energy of each
pick's random 54-atom cell (v3 standard) against v5's recorded prediction and the two ensembles on
the identical cell: Mo₄₂Ta₂₉Ti₂₉ −56.5 (v5 −56.9; v5×3 −50.9 ± 7.7; e9×3 −69.5 ± 1.2);
Mo₇₀Hf₂₀Ti₁₀ −31.2 (−29.6; −24.9 ± 6.0; −33.5 ± 8.5); Mo₄₁Ti₂₅Ta₁₀Nb₈W₇Hf₄Zr₄ −14.3 (−29.8;
−21.3 ± 8.0; −27.2 ± 1.8). Mean |error|: v5 5.8, v5×3 6.3, e9×3 9.4 — and confidently wrong:
−13 on two of three with a seed spread of 1.2–1.8. E229's two faces, now on the search's own
targets: embedding 9 fixes ordered states and over-binds random solutions of unseen compositions,
and its seed spread does not flag it. Rung 0 (random-solution energy) and rung 1 (ordering) now
favour different models; the model that ships must pass both — E237's middle embeddings and E233's
10 Å pairs are the candidates. The random-cell energies of the three picks are confirmed at rung 4
by v5 to within 2 meV on two of three; their ORDERING remains the open question.
Rung-1 speed SOLVED (14:0x) — the exact evaluator.forager/ladder/ece_distill.py::ExactECE
(Fable-written) re-evaluates pyeCE's forward exactly in numpy: per-orbit correlation tensors, the
linear map folded into the first network layer, no per-cluster-function gather. Acceptance on the
teacher runs/ece_v5e4_seed0, Mo₄₂Ta₂₉Ti₂₉, the production cell (432 sites) and settings
(scripts/validate/distill_check.py, runs/distill_check/): per-swap energy differences against
pyeCE's own forward MAE 0.0012 meV, max 0.0071 meV over 2,304 swaps whose mean size is
160–350 meV (float32 rounding), on configurations sampled by either Hamiltonian at 300–1500 K;
rung 1 on the same seed 1148 ± 114 K against the teacher's 1175 ± 114 K (the difference
is chain divergence from rounding, inside one grid step's σ). Throughput, swap_delta moves/s, one
thread, 432 sites: v5 2,788 (pyeCE forward 1,009), embedding 4 1,987 (403), embedding 9
1,519 (28 — 54×). Rung-1 cost no longer depends materially on embedding size, so E237's
decision reduces to accuracy: prediction (3) is moot, and the embedding is chosen by the DFT
targets alone. The fitted alternative (DistilledCE, pairs + triplets distilled from the teacher)
FAILS parity (fresh-state swap RMS 55.7 meV) although its T_c landed at 1170 K — a surrogate can
hit one number by accident, which is why parity, not T_c, is the acceptance test; it is kept only as
the recorded negative. Rung 1 in ECEVerifier now samples the MEAN of every ensemble member through
the exact evaluator (_hamiltonian), so rung 0 and rung 1 score the same model; rung1_model
names every member.
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 15451–15489
E223 — the first search walked up the corrected ladder, pre-registered (2026-09-22 04:2x; chain armed, gated on the search slot)
Gap 2: rung 1 became the Metropolis sampler on the physical axis on 2026-09-21 and has since
run only on hand-picked systems (E210b, E190b's re-read, E222's cell). No composition a
SEARCH proposed has ever been through it.runs/e223_chain.sh (referee-written, verified
by running its parts): one free search on E201b's arms and budget (elitist, fly; 200 × 4 × 2
seeds; pure reward; the fb,oa,5ht brain flags) with FORAGER_RUNG0=ece and no --confirm
(rung 1 inline is 20–32 min per qualifier), then scripts/search/walk_finds.py (new) walks
the pooled niche leaders — de-duplicated on composition, capped at 32 — up rung 1 and rung 2
(MACE hull), one core, ~25 min each, each result on disk before the next. Rung 3 is not
walked while trust_kinetics is False. The top 3 by ladder p not already in data/dft are
written as 54-atom cells to runs/e223_dft_queue.txt; pw.x is not launched. Design:
docs/research/2026-09-22-e223-first-search-on-corrected-ladder.md. One deviation from
E201b: ce_12element restricted by FORAGER_ELEMENTS to the eCE's nine, so every leader is
walkable and Cr — where v5 and SOB20 disagree — is in the space.
Two ladder defects found while wiring it, both fixed before the chain was armed. (i)
ECEVerifier was not callable: Verifier.__call__ implements the walk, the wrapper delegates
by __getattr__, and Python looks special methods up on the type — V(x) raised, and
V.base(x) would silently have walked from the MACE-frame screen. ECEVerifier.__call__ now
runs the base walk with the wrapper's own screen. (ii) A censored rung 1 reached rung 2 as
T = NaN and came out as p = NaN — neither pass nor fail; Verifier._requirement now
returns the safe zero with rung1_status at rungs 2 and 3. Both pinned in
tests/test_rung1_metropolis.py. Consequence the referee measured: with a 227 K grid floor
the ladder can pass a composition only on the HIGH side of the window; "below 90 K" is
unreachable at this budget.
Predictions (on the ladder's own convention — cooling leg, largest peak: E210b reads
747 / 463 / 1163 / 1210 K that way, two of four above 1000). (1) Of 32 leaders, 16 ± 6
pass high (finite T > 1000 K), ~11 fail inside the window, ~5 censored; fewer than 6 passing
high falsifies the reading that the ×2 correction moves search finds up as it moved the
scorecard. (2) The fly's leaders do not differ from elitist's in element support (mean
element count within 0.5, both 3.0–3.6; ≥ 60 % of supports reached by both; Cr in at most 4
of 32). (3) Success = at least one composition with finite rung-2 p ≥ 0.8 and drive_cold ≤
−40 meV/atom; falsified if zero of 32 reach finite p > 0.5, or every pass is a censored
NaN, or every pass is a composition data/dft already holds. Wall time: search 2–3 h, walk
~13 h on one core. Verified before arming: the dry run on the exact Ising reaches 1094 K on
Mo₅₀Ta₅₀ against 1409 at a 167 K grid; one real composition went rung 0 + 1 + 2 in 18 s;
random_cell.py on a pick reproduced E222's own numbers.
EXPERIMENTS.md · lines 16007–16036
E223 rung-4 RESULT (14:1x, Colab A100) — the picks' random cells. DFT formation energy of each
pick's random 54-atom cell (v3 standard) against v5's recorded prediction and the two ensembles on
the identical cell: Mo₄₂Ta₂₉Ti₂₉ −56.5 (v5 −56.9; v5×3 −50.9 ± 7.7; e9×3 −69.5 ± 1.2);
Mo₇₀Hf₂₀Ti₁₀ −31.2 (−29.6; −24.9 ± 6.0; −33.5 ± 8.5); Mo₄₁Ti₂₅Ta₁₀Nb₈W₇Hf₄Zr₄ −14.3 (−29.8;
−21.3 ± 8.0; −27.2 ± 1.8). Mean |error|: v5 5.8, v5×3 6.3, e9×3 9.4 — and confidently wrong:
−13 on two of three with a seed spread of 1.2–1.8. E229's two faces, now on the search's own
targets: embedding 9 fixes ordered states and over-binds random solutions of unseen compositions,
and its seed spread does not flag it. Rung 0 (random-solution energy) and rung 1 (ordering) now
favour different models; the model that ships must pass both — E237's middle embeddings and E233's
10 Å pairs are the candidates. The random-cell energies of the three picks are confirmed at rung 4
by v5 to within 2 meV on two of three; their ORDERING remains the open question.
Rung-1 speed SOLVED (14:0x) — the exact evaluator.forager/ladder/ece_distill.py::ExactECE
(Fable-written) re-evaluates pyeCE's forward exactly in numpy: per-orbit correlation tensors, the
linear map folded into the first network layer, no per-cluster-function gather. Acceptance on the
teacher runs/ece_v5e4_seed0, Mo₄₂Ta₂₉Ti₂₉, the production cell (432 sites) and settings
(scripts/validate/distill_check.py, runs/distill_check/): per-swap energy differences against
pyeCE's own forward MAE 0.0012 meV, max 0.0071 meV over 2,304 swaps whose mean size is
160–350 meV (float32 rounding), on configurations sampled by either Hamiltonian at 300–1500 K;
rung 1 on the same seed 1148 ± 114 K against the teacher's 1175 ± 114 K (the difference
is chain divergence from rounding, inside one grid step's σ). Throughput, swap_delta moves/s, one
thread, 432 sites: v5 2,788 (pyeCE forward 1,009), embedding 4 1,987 (403), embedding 9
1,519 (28 — 54×). Rung-1 cost no longer depends materially on embedding size, so E237's
decision reduces to accuracy: prediction (3) is moot, and the embedding is chosen by the DFT
targets alone. The fitted alternative (DistilledCE, pairs + triplets distilled from the teacher)
FAILS parity (fresh-state swap RMS 55.7 meV) although its T_c landed at 1170 K — a surrogate can
hit one number by accident, which is why parity, not T_c, is the acceptance test; it is kept only as
the recorded negative. Rung 1 in ECEVerifier now samples the MEAN of every ensemble member through
the exact evaluator (_hamiltonian), so rung 0 and rung 1 score the same model; rung1_model
names every member.
EXPERIMENTS.md · lines 16353–16360
E223 scored — the walk it pre-registered.runs/e223_walk/summary.jsonl (32 leaders; v5,
old rung-1 floor). (1) Falsified: 1 of 32 passes high (Mo₇₀Hf₂₀Ti₁₀, 1291 K) against 16 ± 6
(fewer than 6 falsifies); 7 order inside the window (424–896 K), 24 show no transition above
100 K. (2) Falsified: the arms differ in support: fly 3.17 elements per leader, elitist 5.21
(elements ≥ 1 at.%), and 1 of 21 supports is reached by both (≥ 60 % predicted); Cr is in 4 of
32 (at most 4: holds). (3) Neither branch: best finite p 0.572 (Mo₄₂Ta₂₉Ti₂₉, drive_cold
−114), below the 0.8 success bar and above the 0.5 falsifier. Superseded as a verdict on the
finds by E236/E240, which re-walked the same 32 on the corrected ladder (e6, floor 60 K).
Related entries
E210b — the nine-system scorecard by the second sampler, pre-registered (2026-09-21 19:1x)