Does letting atom pairs interact over a longer distance make the energy model more accurate?
Partly. Error on 11 DFT test cells fell from 20.6 to 15.9 meV/atom, just missing the bar of 15; the Mo–Ta miss did not shrink.
In the log: the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)
mixedDate 2026-09-22 20:2x, as written in the logrung 1 · ordering3 predictions · 1 result paragraphEXPERIMENTS.md lines 15852–15860, lines 15979–16005
What E233 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E233.svg).
Pre-registration
(1)
Ensemble MAE on the eleven
cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2.
no verdict written against it
(2)
Ensemble-mean training RMSE from 13.6 to ≤ 12.1.
confirmed
(3)
Falsifier: eleven-cell MAE
within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234)
carry the programme. Chain runs/e233_chain.sh, after E231.
no verdict written against it
The pre-registration, as written
E233 — the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)
One variable: --cutoffs pair max 6.0 → 10.0 Å (15 pair orbits instead of 5; triplets 4.5 Å
unchanged), v5's labels, rows, recipe and seeds 0–2 → ece_v5c10_seed*; scored as an ensemble
on the unbiased set against the three v5 seeds. Predictions. (1) Ensemble MAE on the eleven
cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2.
(2) Ensemble-mean training RMSE from 13.6 to ≤ 12.1. (3) Falsifier: eleven-cell MAE
within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234)
carry the programme. Chain runs/e233_chain.sh, after E231.
Results
EXPERIMENTS.md · line 15979
E233 RESULT (12:5x) — pair range helps everything except Mo–Ta. Pairs to 10 Å, v5's recipe,
3 seeds (GPU): eleven-cell MAE 15.9 (v5 20.6), validation RMSE 13.3 (17.7), training
10.5 (13.6), held-48 6.5 (7.1); B2 Mo–Ta +28.5 (+29.5), four Mo–Ta cells' mean +30.5
(+34). (1) MAE ≤ 15 missed by 0.9, Mo–Ta < +20 missed; held-48 met. (2) confirmed. (3) not
triggered (MAE moved 4.7). Range improves generalisation; only embedding size moves Mo–Ta.
Rung-1 cost, measured (scripts/validate/rung1_speed.py: FastECE.swap_delta moves/s, 432
sites, one thread, Mo₄₂Ta₂₉Ti₂₉): v5 1009, embedding 4 403, pairs 10 Å 158, embedding 9
28. A v5 rung-1 walk is ~25 min, so ~1 h / ~2.7 h / ~15 h respectively. Accuracy and rung-1 cost
now pull apart; if E237's smallest adequate embedding is too slow once combined with 10 Å pairs, the
answer is a fast rung-1 surrogate distilled from the accurate model (Fable-written, verified
against direct Monte Carlo) rather than a compromise model.
Data bug found (13:3x): 18 training rows were written as the wrong decoration in every eCE fit.train_ece.py writes a RHEA row's symbols in file order onto ideal_frac_of(n)'s site order,
assuming RHEA stores atoms in that order. Checked on all 4,326 of v5's rows against each frame's
own ideal-site mapping: true for 4,308, false for 18 — one sub-batch of bcc_alloys_ordered
(indices 8132–8152, 8248–8256: 11 × 16-atom, 6 × 54-atom, 1 × 128-atom; Cr–V, V–W, Ta–W, Cr–W,
Cr–Ta orderings at −24.6 … +10.5 meV/atom), listed in runs/scrambled_rows.json. Each taught the
model an ordered-state energy on a scrambled, near-random decoration. 0.4 % of rows and not the
deep ones (B2 Mo–Ta, index 6973, is correct), so it cannot be the +30 Mo–Ta error, but it biases
ordering energetics in five binaries. Every comparison made today (E229, E231, E233, E237) carries
the same 18 rows in every arm, so relative conclusions stand; the model that ships is refitted on
corrected emission. E234 FAILED on the same path: emit silently skips any row whose n is
not a cubic 2m³ count (if len(frac) != r["n"]: continue), so pyeCE received 4,189 structures
while the split counted ~5,880 — it crashed (index 4190 of 4189) instead of misaligning silently.
Fix delegated (Fable, isolated worktree): rows carry their own ideal-site positions_frac (and
cell when non-cubic); emit never skips silently; the mapped count must equal the written count.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 15852–15860
E233 — the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)
One variable: --cutoffs pair max 6.0 → 10.0 Å (15 pair orbits instead of 5; triplets 4.5 Å
unchanged), v5's labels, rows, recipe and seeds 0–2 → ece_v5c10_seed*; scored as an ensemble
on the unbiased set against the three v5 seeds. Predictions. (1) Ensemble MAE on the eleven
cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2.
(2) Ensemble-mean training RMSE from 13.6 to ≤ 12.1. (3) Falsifier: eleven-cell MAE
within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234)
carry the programme. Chain runs/e233_chain.sh, after E231.
EXPERIMENTS.md · lines 15979–16005
E233 RESULT (12:5x) — pair range helps everything except Mo–Ta. Pairs to 10 Å, v5's recipe,
3 seeds (GPU): eleven-cell MAE 15.9 (v5 20.6), validation RMSE 13.3 (17.7), training
10.5 (13.6), held-48 6.5 (7.1); B2 Mo–Ta +28.5 (+29.5), four Mo–Ta cells' mean +30.5
(+34). (1) MAE ≤ 15 missed by 0.9, Mo–Ta < +20 missed; held-48 met. (2) confirmed. (3) not
triggered (MAE moved 4.7). Range improves generalisation; only embedding size moves Mo–Ta.
Rung-1 cost, measured (scripts/validate/rung1_speed.py: FastECE.swap_delta moves/s, 432
sites, one thread, Mo₄₂Ta₂₉Ti₂₉): v5 1009, embedding 4 403, pairs 10 Å 158, embedding 9
28. A v5 rung-1 walk is ~25 min, so ~1 h / ~2.7 h / ~15 h respectively. Accuracy and rung-1 cost
now pull apart; if E237's smallest adequate embedding is too slow once combined with 10 Å pairs, the
answer is a fast rung-1 surrogate distilled from the accurate model (Fable-written, verified
against direct Monte Carlo) rather than a compromise model.
Data bug found (13:3x): 18 training rows were written as the wrong decoration in every eCE fit.train_ece.py writes a RHEA row's symbols in file order onto ideal_frac_of(n)'s site order,
assuming RHEA stores atoms in that order. Checked on all 4,326 of v5's rows against each frame's
own ideal-site mapping: true for 4,308, false for 18 — one sub-batch of bcc_alloys_ordered
(indices 8132–8152, 8248–8256: 11 × 16-atom, 6 × 54-atom, 1 × 128-atom; Cr–V, V–W, Ta–W, Cr–W,
Cr–Ta orderings at −24.6 … +10.5 meV/atom), listed in runs/scrambled_rows.json. Each taught the
model an ordered-state energy on a scrambled, near-random decoration. 0.4 % of rows and not the
deep ones (B2 Mo–Ta, index 6973, is correct), so it cannot be the +30 Mo–Ta error, but it biases
ordering energetics in five binaries. Every comparison made today (E229, E231, E233, E237) carries
the same 18 rows in every arm, so relative conclusions stand; the model that ships is refitted on
corrected emission. E234 FAILED on the same path: emit silently skips any row whose n is
not a cubic 2m³ count (if len(frac) != r["n"]: continue), so pyeCE received 4,189 structures
while the split counted ~5,880 — it crashed (index 4190 of 4189) instead of misaligning silently.
Fix delegated (Fable, isolated worktree): rows carry their own ideal-site positions_frac (and
cell when non-cubic); emit never skips silently; the mapped count must equal the written count.
Related entries
E231 — the E228 relabel, one variable: v5's rows, v5's recipe, corrected labels (pre-registered…
E234 — the ordered cells the labeller dropped (pre-registered 2026-09-22 20:3x)
E237 — the smallest embedding that holds Mo–Ta (pre-registered 2026-09-23 12:3x)
E229 — is it capacity? v5's recipe with one embedding dimension per species (pre-registered…