Experiments · E233

Does letting atom pairs interact over a longer distance make the energy model more accurate?

Partly. Error on 11 DFT test cells fell from 20.6 to 15.9 meV/atom, just missing the bar of 15; the Mo–Ta miss did not shrink.

In the log: the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)

mixedDate 2026-09-22 20:2x, as written in the logrung 1 · ordering3 predictions · 1 result paragraphEXPERIMENTS.md lines 15852–15860, lines 15979–16005
exp E233 diagram
What E233 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E233.svg).

Pre-registration

  1. (1)
    Ensemble MAE on the eleven cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2.
    no verdict written against it
  2. (2)
    Ensemble-mean training RMSE from 13.6 to ≤ 12.1.
    confirmed
  3. (3)
    Falsifier: eleven-cell MAE within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234) carry the programme. Chain runs/e233_chain.sh, after E231.
    no verdict written against it
The pre-registration, as written

E233 — the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)

One variable: --cutoffs pair max 6.0 → 10.0 Å (15 pair orbits instead of 5; triplets 4.5 Å unchanged), v5's labels, rows, recipe and seeds 0–2 → ece_v5c10_seed*; scored as an ensemble on the unbiased set against the three v5 seeds. Predictions. (1) Ensemble MAE on the eleven cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2. (2) Ensemble-mean training RMSE from 13.6 to ≤ 12.1. (3) Falsifier: eleven-cell MAE within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234) carry the programme. Chain runs/e233_chain.sh, after E231.

Results

EXPERIMENTS.md · line 15979

E233 RESULT (12:5x) — pair range helps everything except Mo–Ta. Pairs to 10 Å, v5's recipe, 3 seeds (GPU): eleven-cell MAE 15.9 (v5 20.6), validation RMSE 13.3 (17.7), training 10.5 (13.6), held-48 6.5 (7.1); B2 Mo–Ta +28.5 (+29.5), four Mo–Ta cells' mean +30.5 (+34). (1) MAE ≤ 15 missed by 0.9, Mo–Ta < +20 missed; held-48 met. (2) confirmed. (3) not triggered (MAE moved 4.7). Range improves generalisation; only embedding size moves Mo–Ta. Rung-1 cost, measured (scripts/validate/rung1_speed.py: FastECE.swap_delta moves/s, 432 sites, one thread, Mo₄₂Ta₂₉Ti₂₉): v5 1009, embedding 4 403, pairs 10 Å 158, embedding 9 28. A v5 rung-1 walk is ~25 min, so ~1 h / ~2.7 h / ~15 h respectively. Accuracy and rung-1 cost now pull apart; if E237's smallest adequate embedding is too slow once combined with 10 Å pairs, the answer is a fast rung-1 surrogate distilled from the accurate model (Fable-written, verified against direct Monte Carlo) rather than a compromise model.

Data bug found (13:3x): 18 training rows were written as the wrong decoration in every eCE fit. train_ece.py writes a RHEA row's symbols in file order onto ideal_frac_of(n)'s site order, assuming RHEA stores atoms in that order. Checked on all 4,326 of v5's rows against each frame's own ideal-site mapping: true for 4,308, false for 18 — one sub-batch of bcc_alloys_ordered (indices 8132–8152, 8248–8256: 11 × 16-atom, 6 × 54-atom, 1 × 128-atom; Cr–V, V–W, Ta–W, Cr–W, Cr–Ta orderings at −24.6 … +10.5 meV/atom), listed in runs/scrambled_rows.json. Each taught the model an ordered-state energy on a scrambled, near-random decoration. 0.4 % of rows and not the deep ones (B2 Mo–Ta, index 6973, is correct), so it cannot be the +30 Mo–Ta error, but it biases ordering energetics in five binaries. Every comparison made today (E229, E231, E233, E237) carries the same 18 rows in every arm, so relative conclusions stand; the model that ships is refitted on corrected emission. E234 FAILED on the same path: emit silently skips any row whose n is not a cubic 2m³ count (if len(frac) != r["n"]: continue), so pyeCE received 4,189 structures while the split counted ~5,880 — it crashed (index 4190 of 4189) instead of misaligning silently. Fix delegated (Fable, isolated worktree): rows carry their own ideal-site positions_frac (and cell when non-cubic); emit never skips silently; the mapped count must equal the written count.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 15852–15860

E233 — the pair range: v5's recipe with pairs to 10 Å (pre-registered 2026-09-22 20:2x)

One variable: --cutoffs pair max 6.0 → 10.0 Å (15 pair orbits instead of 5; triplets 4.5 Å unchanged), v5's labels, rows, recipe and seeds 0–2 → ece_v5c10_seed*; scored as an ensemble on the unbiased set against the three v5 seeds. Predictions. (1) Ensemble MAE on the eleven cells from 20.6 to ≤ 15, the four Mo–Ta cells' mean from +34 to < +20, held-48 ≤ 8.2. (2) Ensemble-mean training RMSE from 13.6 to ≤ 12.1. (3) Falsifier: eleven-cell MAE within ±3 of 20.6 → pair range is not the constraint either, and the data levers (E231, E234) carry the programme. Chain runs/e233_chain.sh, after E231.

EXPERIMENTS.md · lines 15979–16005

E233 RESULT (12:5x) — pair range helps everything except Mo–Ta. Pairs to 10 Å, v5's recipe, 3 seeds (GPU): eleven-cell MAE 15.9 (v5 20.6), validation RMSE 13.3 (17.7), training 10.5 (13.6), held-48 6.5 (7.1); B2 Mo–Ta +28.5 (+29.5), four Mo–Ta cells' mean +30.5 (+34). (1) MAE ≤ 15 missed by 0.9, Mo–Ta < +20 missed; held-48 met. (2) confirmed. (3) not triggered (MAE moved 4.7). Range improves generalisation; only embedding size moves Mo–Ta. Rung-1 cost, measured (scripts/validate/rung1_speed.py: FastECE.swap_delta moves/s, 432 sites, one thread, Mo₄₂Ta₂₉Ti₂₉): v5 1009, embedding 4 403, pairs 10 Å 158, embedding 9 28. A v5 rung-1 walk is ~25 min, so ~1 h / ~2.7 h / ~15 h respectively. Accuracy and rung-1 cost now pull apart; if E237's smallest adequate embedding is too slow once combined with 10 Å pairs, the answer is a fast rung-1 surrogate distilled from the accurate model (Fable-written, verified against direct Monte Carlo) rather than a compromise model.

Data bug found (13:3x): 18 training rows were written as the wrong decoration in every eCE fit. train_ece.py writes a RHEA row's symbols in file order onto ideal_frac_of(n)'s site order, assuming RHEA stores atoms in that order. Checked on all 4,326 of v5's rows against each frame's own ideal-site mapping: true for 4,308, false for 18 — one sub-batch of bcc_alloys_ordered (indices 8132–8152, 8248–8256: 11 × 16-atom, 6 × 54-atom, 1 × 128-atom; Cr–V, V–W, Ta–W, Cr–W, Cr–Ta orderings at −24.6 … +10.5 meV/atom), listed in runs/scrambled_rows.json. Each taught the model an ordered-state energy on a scrambled, near-random decoration. 0.4 % of rows and not the deep ones (B2 Mo–Ta, index 6973, is correct), so it cannot be the +30 Mo–Ta error, but it biases ordering energetics in five binaries. Every comparison made today (E229, E231, E233, E237) carries the same 18 rows in every arm, so relative conclusions stand; the model that ships is refitted on corrected emission. E234 FAILED on the same path: emit silently skips any row whose n is not a cubic 2m³ count (if len(frac) != r["n"]: continue), so pyeCE received 4,189 structures while the split counted ~5,880 — it crashed (index 4190 of 4189) instead of misaligning silently. Fix delegated (Fable, isolated worktree): rows carry their own ideal-site positions_frac (and cell when non-cubic); emit never skips silently; the mapped count must equal the written count.

Related entries

Built with PRISMWebsite and visualizations made using Claude