Experiments · E210b

Does the corrected sampler reproduce published ordering temperatures across nine alloys?

Partly. Six comparable systems land within 22 % of the published values; Cr-Ta-Ti-W is off by a factor of 2.92.

In the log: the nine-system scorecard by the second sampler, pre-registered (2026-09-21 19:1x)

mixedDate 2026-09-21 19:1x, as written in the logrung 4 · DFT5 predictions · 0 result paragraphsEXPERIMENTS.md lines 14895–14927, lines 14948–14966, lines 15103–15114, lines 15344–15361, lines 15388–15391, lines 15651–15668, lines 16337–16351
exp E210b diagram
What E210b did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E210b.svg).

Pre-registration

  1. (1)
    every system is reversible: floor mismatch < 3 meV, the two legs' dE/dT peaks within one grid step — a system that is not has its number withheld.
    no verdict written against it
  2. (2)
    v5's dE/dT T_c is 0.5–1.0× the like-for-like ideal-lattice rows (KOS19 MoNbTaW 600, KW23 MoNbTaW 1110 and Mo–Ta 2020): the measured points so far are 0.88× (MoNbTaW, censored ≥ 528), ~0.5× (Mo–Ta), 0.93× (MoNbTaVW against an H_mix row). E210's original "0.3–0.6×" was on the nominal axis.
    no verdict written against it
  3. (3)
    no SRO-onset observable is claimed (ledger, 2026-09-21); SRO is reported as decoded curves and compared where SOB20/FC17 publish H_mix inflections only as context.
    no verdict written against it
  4. (4)
    the V-bearing (B32-like) systems show a lower v5/literature ratio than the valence-difference B2 systems.
    no verdict written against it
  5. (5)
    cross-check: on Cr20Ta20Ti20V20W20 the pyeCE run's E(T_nominal) equals this sampler's E(2·T_nominal) within 1 meV/atom over the overlap 200–2600 K physical, as E214 did for Mo–Ta (0.41 meV). Scoring: published_odt.comparable() per row; results in runs/e210b_scorecard/<system>/mc.json.
    no verdict written against it
The pre-registration, as written

E210b — the nine-system scorecard by the second sampler, pre-registered (2026-09-21 19:1x)

Operator's decision after the Colab question: the pyeCE E210 run is stopped after its first system (Cr20Ta20Ti20V20W20 finishes as a cross-check; the chain that would launch the other eight is killed) and the scorecard is run by the second sampler instead — locally, one core, runs/e210b_chain.sh. Colab was declined on the numbers: a sequential Metropolis chain gains nothing from a GPU, Colab's CPU is ≥ 8× slower than this machine for exactly this loop, the CLI handle cannot hold a multi-hour job, and it would ship the trained model off the machine.

The instrument. scripts/ordering/v5_metropolis.py, generalised today: any composition in the PRIM's nine elements (--alloy Cr20Ta20Ti20V20W20, largest-remainder counts on 432 sites), Warren–Cowley α for every unlike pair on four shells (alpha{k}_pairs; alpha{k} is the strongest ordering pair, the convention the pyeCE reader uses), a partial_<leg>.json after every temperature, done only after both legs. Tested against forager.physics.order.warren_cowley on a B2-like (Mo,W | Nb,Ta) cell to 1e-9 for every pair (tests/test_v5_metropolis_multicomponent.py). It samples at the temperature it reports and needs no decoding — the two defects of the pyeCE path are absent by construction. Protocol: 2600 → 100 → 2600 K physical, 35 rungs per leg (73.5 K), 150 + 350 sweeps (E216's), ~2 h per system on one core.

Predictions, carried over from E210 (2026-09-20) and restated on the physical axis. (1) every system is reversible: floor mismatch < 3 meV, the two legs' dE/dT peaks within one grid step — a system that is not has its number withheld. (2) v5's dE/dT T_c is 0.5–1.0× the like-for-like ideal-lattice rows (KOS19 MoNbTaW 600, KW23 MoNbTaW 1110 and Mo–Ta 2020): the measured points so far are 0.88× (MoNbTaW, censored ≥ 528), ~0.5× (Mo–Ta), 0.93× (MoNbTaVW against an H_mix row). E210's original "0.3–0.6×" was on the nominal axis. (3) no SRO-onset observable is claimed (ledger, 2026-09-21); SRO is reported as decoded curves and compared where SOB20/FC17 publish H_mix inflections only as context. (4) the V-bearing (B32-like) systems show a lower v5/literature ratio than the valence-difference B2 systems. (5) cross-check: on Cr20Ta20Ti20V20W20 the pyeCE run's E(T_nominal) equals this sampler's E(2·T_nominal) within 1 meV/atom over the overlap 200–2600 K physical, as E214 did for Mo–Ta (0.41 meV). Scoring: published_odt.comparable() per row; results in runs/e210b_scorecard/<system>/mc.json.

E210b, first system (2026-09-21 19:5x): Cr20Ta20Ti20V20W20. 6³, 35 + 35 rungs (73.5 K), 150 + 350 sweeps, 3.1 h on one core. dE/dT peak 747 ± 74 K cooling / 830 ± 184 K heating, both inside the sweep — the 2600 K ceiling does not censor it despite α(Cr–Ta) = −0.15 already at the ceiling; variance-channel peaks 693 / 475 (the heating one weak at 350 sweeps). Floor order: Ta–Ti is the strongest pair (α₁ −0.47 / −0.49), then V–W (−0.31), Cr–W (−0.27), Cr–Ta (−0.25); E from +46.1 meV/atom at 2600 K to −40.1 at 100 K; junction mismatch −0.06 meV. Scored: (1) reversible on energy (< 3 meV) — the two dE/dT peaks are 83 K apart against a 73.5 K grid step, ten kelvin past the pre-registered "within one step", and inside the heating peak's own 184 K width; reported as unresolved, not as hysteresis. (2) no cv_peak/chi_peak row exists for this system; against SOB20's F_mix crossing 900 K and H_mix inflection 1000 K (context only, not like-for-like) v5 sits at 0.75–0.92×. (5) cross-check against the pyeCE run's cooling leg (scripts/ordering/e210b_crosscheck.py): E_pyeCE(T_nominal) − E_sampler(2 T) has mean |diff| 1.05 meV/atom over 17 overlapping states (bar 1.0) — 0.3–0.9 meV away from the transition and 2.9–3.8 meV on the two rungs straddling it (935 and 788 K physical), where 350 sweeps of two different samplers disagree on how far the order has run. The factor of two holds on a quinary as it did on Mo–Ta; the bar is missed by 0.05 meV because of the transition rungs alone. The α columns of that table compare different pairs (the pyeCE reader picks V–W as its strongest decoded pair, the sampler Ta–Ti) and are not a like-for-like check. Eight systems remain, ~3 h each.

E210b, second system (2026-09-22 00:2x): Ta25Ti25V25W25. dE/dT peak 463 ± 74 K cooling / 546 ± 74 K heating (variance 471 / 525), both inside the sweep; junction mismatch −0.07 meV; floor order Ta–W (α₁ −0.47), V–W (−0.40), Ta–Ti (−0.33), with Ta–V (+0.39) and Ti–V (+0.28) avoiding; E from −2 meV/atom at 2600 K to −66 at 100 K. Against SOB20's F_mix crossing at 500 K (context, not like-for-like) the two legs bracket it: 0.93× / 1.09×. Second system in a row where heating sits ~80 K above cooling — one rung plus ten kelvin — with the energy reversible to < 0.1 meV: at 150 + 350 sweeps the cooling leg undercools and the heating leg overshoots by about half a rung each, so the pair brackets the transition and their midpoint (505 K here, 789 K for Cr20Ta20Ti20V20W20) is the estimate, the half-gap (±40 K) its bar. The pre-registered "within one grid step" is missed by the same ten kelvin twice; the protocol is reversible on energy and the gap is a sweep-count effect, not hysteresis. Third system (Cr25Ta25Ti25W25) running.

E210b, fourth system, and the scorecard so far (2026-09-22 02:5x). Cr₂₅Ta₂₅V₂₅W₂₅: highest-T peak 1272 K (1210/1335) against SOB20's ODTT 1300 — ratio 0.98. The four finished systems, scored like-for-like against the corrected table:

system largest highest-T SOB20 ODTT ratio floor α₁
Cr20Ta20Ti20V20W20 788 1076 1000 1.08 −0.47 (Ta–Ti)
Ta25Ti25V25W25 505 607 500 1.21 −0.47 (Ta–W)
Cr25Ta25V25W25 1031 1272 1300 0.98 −0.45 (Cr–Ta)
Cr25Ta25Ti25W25 1198 1461 500 2.92 −1.03 (Cr–Ta)

Three of four within 21 %, one of them to 2 %. This is the rung that was "0.5× published and under active-learning repair" on 2026-09-18; the repair was a sampler that ran at twice its stated temperature and a table that compared a heat-capacity peak with a free-energy crossing. Note what the fourth column does not say: v5 puts its strongest ordering on Cr–Ta in both Cr-bearing systems, where SOB20 report Cr–V. On Cr–Ta–V–W the temperature still agrees to 2 % with the wrong pair driving it, which is a warning about reading agreement as validation — and precisely why E222 tests the pair and not the temperature.

E210b, fifth system: Cr₂₅Ti₂₅V₂₅W₂₅ — largest peak 796 / 947 K (cooling/heating), highest-T 796 / 1151 (the heating leg carries a small hot bump), floor α₁ −0.34 (Cr–Ti), E −54 meV/atom, reversible. Not scoreable yet: the table holds only its F_mix value (700 K); the H_mix ODTT is being retrieved from SOB20's Table 2. Note the strongest pair here is Cr–Ti, not Cr–Ta or Cr–V.

E210b, systems 6–8. Cr₂₅Ta₂₅Ti₂₅V₂₅: highest-T 646 K (563/730) vs SOB20 700 → 0.92, pairs Cr–Ta −0.64 and Cr–V −0.55 (SOB20: Cr–V). Mo₂₀Nb₂₀Ta₂₀V₂₀W₂₀: 841 K (916/765; largest 733) vs FC17's 750 H_mix inflection → 1.12 (0.98 by the largest peak), pairs Mo–Ta −1.49, Mo–Nb −1.00. Mo₃₃Nb₃₃Ta₃₃: 1100 K, no resolved published value (WAK23), pairs Mo–Ta −0.87. All three reversible (junction ≤ 0.02 meV). Scorecard so far: 1.08, 1.21, 0.98, 1.22, 0.92, 1.12 — six of six comparable systems within 22 %, plus the Cr–Ta–Ti–W 2.92 that E222 explains. The ninth (MoNbTaW, the only cv_peak/chi_peak row) is running.

Seed variance, two of three seeds. v5 refitted with seed 1: held-out MAE 7.4 against seed 0's 7.2 (RMSE 10.7 vs 9.2, bias +2.6 vs +0.5). Preliminary spread ≈ 0.2 meV in MAE, so v7's 9.3 and v7b's 9.9 sit well outside it and their degradation is real, not fit noise — the pre-registered branch "a VASP-RHEA vs QE frame offset must be fitted per group before rows are merged" is the one that applies, pending seed 2 (paused by the guardian at load 13.6). E221 (the kinetics anchor) is on its second cell. The fcc reference redo landed: Fe fcc (non-spin) now fits inside its low bracket at a₀ = 3.4617 Å, E_min −4479.7572 eV/atom (the first pass had read the default 0.96–1.02 bracket, whose minimum sat below 0.96; the _lo_e80 scan at 0.93–0.99 brackets it). All five fcc references (Al, Co, Cu, Fe, Ni) are bracketed; E226 can read them.

E210b, ninth system: MoNbTaW (Mo₂₅Nb₂₅Ta₂₅W₂₅) — v5 orders it at under half the published temperature. runs/e210b_scorecard/Mo25Nb25Ta25W25/mc.json: dE/dT peaks 247 K cooling / 321 K heating (one grid step, 73.5 K, apart); heat-capacity (variance) peaks 174 / 321 K; floors agree to 0.0004 meV/atom, so it is reversible by (1). Strongest pair Mo–Ta, α₁ at the cooling peak −1.38 (B2 on (Mo,W | Nb,Ta)). This is the scorecard's one heat-capacity-peak row and the only like-for-like comparison with the KOS19 / KW23 CE+MC work on this chemistry: 0.41× KOS19's ideal-lattice 600 K by the same observable (variance-peak mean 247 K), 0.47× by the ladder's highest-T convention (284 K), and 0.22–0.26× KW23's 1110 K susceptibility peak. Prediction (2) (0.5–1.0×) fails for MoNbTaW. The scorecard therefore reads: six of seven comparable systems within 22 % (1.08, 1.21, 0.98, 1.22, 0.92, 1.12), MoNbTaW at 0.41, and Cr–Ta–Ti–W at 2.92 (E222). The two rows that fall short are the two from CE + Monte Carlo work on Mo–Nb–Ta–W (KOS19/KW23 here; Mo–Ta at 0.49× v5 and 0.57× e6 in E242 (4)). v5 under-orders the Mo-group B2 systems while matching the SOB20 rows. The survivor is 52 % Mo. This is why E242's DFT, not rung 1, decides it. (KOS19's relaxed-lattice 250–350 K, a different ground state, is not the comparable row and is not scored.)

Results

No result paragraph for this entry was found in the log.

The full record

This entry is written in 7 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 14895–14927

E210b — the nine-system scorecard by the second sampler, pre-registered (2026-09-21 19:1x)

Operator's decision after the Colab question: the pyeCE E210 run is stopped after its first system (Cr20Ta20Ti20V20W20 finishes as a cross-check; the chain that would launch the other eight is killed) and the scorecard is run by the second sampler instead — locally, one core, runs/e210b_chain.sh. Colab was declined on the numbers: a sequential Metropolis chain gains nothing from a GPU, Colab's CPU is ≥ 8× slower than this machine for exactly this loop, the CLI handle cannot hold a multi-hour job, and it would ship the trained model off the machine.

The instrument. scripts/ordering/v5_metropolis.py, generalised today: any composition in the PRIM's nine elements (--alloy Cr20Ta20Ti20V20W20, largest-remainder counts on 432 sites), Warren–Cowley α for every unlike pair on four shells (alpha{k}_pairs; alpha{k} is the strongest ordering pair, the convention the pyeCE reader uses), a partial_<leg>.json after every temperature, done only after both legs. Tested against forager.physics.order.warren_cowley on a B2-like (Mo,W | Nb,Ta) cell to 1e-9 for every pair (tests/test_v5_metropolis_multicomponent.py). It samples at the temperature it reports and needs no decoding — the two defects of the pyeCE path are absent by construction. Protocol: 2600 → 100 → 2600 K physical, 35 rungs per leg (73.5 K), 150 + 350 sweeps (E216's), ~2 h per system on one core.

Predictions, carried over from E210 (2026-09-20) and restated on the physical axis. (1) every system is reversible: floor mismatch < 3 meV, the two legs' dE/dT peaks within one grid step — a system that is not has its number withheld. (2) v5's dE/dT T_c is 0.5–1.0× the like-for-like ideal-lattice rows (KOS19 MoNbTaW 600, KW23 MoNbTaW 1110 and Mo–Ta 2020): the measured points so far are 0.88× (MoNbTaW, censored ≥ 528), ~0.5× (Mo–Ta), 0.93× (MoNbTaVW against an H_mix row). E210's original "0.3–0.6×" was on the nominal axis. (3) no SRO-onset observable is claimed (ledger, 2026-09-21); SRO is reported as decoded curves and compared where SOB20/FC17 publish H_mix inflections only as context. (4) the V-bearing (B32-like) systems show a lower v5/literature ratio than the valence-difference B2 systems. (5) cross-check: on Cr20Ta20Ti20V20W20 the pyeCE run's E(T_nominal) equals this sampler's E(2·T_nominal) within 1 meV/atom over the overlap 200–2600 K physical, as E214 did for Mo–Ta (0.41 meV). Scoring: published_odt.comparable() per row; results in runs/e210b_scorecard/<system>/mc.json.

EXPERIMENTS.md · lines 14948–14966

E210b, first system (2026-09-21 19:5x): Cr20Ta20Ti20V20W20. 6³, 35 + 35 rungs (73.5 K), 150 + 350 sweeps, 3.1 h on one core. dE/dT peak 747 ± 74 K cooling / 830 ± 184 K heating, both inside the sweep — the 2600 K ceiling does not censor it despite α(Cr–Ta) = −0.15 already at the ceiling; variance-channel peaks 693 / 475 (the heating one weak at 350 sweeps). Floor order: Ta–Ti is the strongest pair (α₁ −0.47 / −0.49), then V–W (−0.31), Cr–W (−0.27), Cr–Ta (−0.25); E from +46.1 meV/atom at 2600 K to −40.1 at 100 K; junction mismatch −0.06 meV. Scored: (1) reversible on energy (< 3 meV) — the two dE/dT peaks are 83 K apart against a 73.5 K grid step, ten kelvin past the pre-registered "within one step", and inside the heating peak's own 184 K width; reported as unresolved, not as hysteresis. (2) no cv_peak/chi_peak row exists for this system; against SOB20's F_mix crossing 900 K and H_mix inflection 1000 K (context only, not like-for-like) v5 sits at 0.75–0.92×. (5) cross-check against the pyeCE run's cooling leg (scripts/ordering/e210b_crosscheck.py): E_pyeCE(T_nominal) − E_sampler(2 T) has mean |diff| 1.05 meV/atom over 17 overlapping states (bar 1.0) — 0.3–0.9 meV away from the transition and 2.9–3.8 meV on the two rungs straddling it (935 and 788 K physical), where 350 sweeps of two different samplers disagree on how far the order has run. The factor of two holds on a quinary as it did on Mo–Ta; the bar is missed by 0.05 meV because of the transition rungs alone. The α columns of that table compare different pairs (the pyeCE reader picks V–W as its strongest decoded pair, the sampler Ta–Ti) and are not a like-for-like check. Eight systems remain, ~3 h each.

EXPERIMENTS.md · lines 15103–15114

E210b, second system (2026-09-22 00:2x): Ta25Ti25V25W25. dE/dT peak 463 ± 74 K cooling / 546 ± 74 K heating (variance 471 / 525), both inside the sweep; junction mismatch −0.07 meV; floor order Ta–W (α₁ −0.47), V–W (−0.40), Ta–Ti (−0.33), with Ta–V (+0.39) and Ti–V (+0.28) avoiding; E from −2 meV/atom at 2600 K to −66 at 100 K. Against SOB20's F_mix crossing at 500 K (context, not like-for-like) the two legs bracket it: 0.93× / 1.09×. Second system in a row where heating sits ~80 K above cooling — one rung plus ten kelvin — with the energy reversible to < 0.1 meV: at 150 + 350 sweeps the cooling leg undercools and the heating leg overshoots by about half a rung each, so the pair brackets the transition and their midpoint (505 K here, 789 K for Cr20Ta20Ti20V20W20) is the estimate, the half-gap (±40 K) its bar. The pre-registered "within one grid step" is missed by the same ten kelvin twice; the protocol is reversible on energy and the gap is a sweep-count effect, not hysteresis. Third system (Cr25Ta25Ti25W25) running.

EXPERIMENTS.md · lines 15344–15361

E210b, fourth system, and the scorecard so far (2026-09-22 02:5x). Cr₂₅Ta₂₅V₂₅W₂₅: highest-T peak 1272 K (1210/1335) against SOB20's ODTT 1300 — ratio 0.98. The four finished systems, scored like-for-like against the corrected table:

system largest highest-T SOB20 ODTT ratio floor α₁
Cr20Ta20Ti20V20W20 788 1076 1000 1.08 −0.47 (Ta–Ti)
Ta25Ti25V25W25 505 607 500 1.21 −0.47 (Ta–W)
Cr25Ta25V25W25 1031 1272 1300 0.98 −0.45 (Cr–Ta)
Cr25Ta25Ti25W25 1198 1461 500 2.92 −1.03 (Cr–Ta)

Three of four within 21 %, one of them to 2 %. This is the rung that was "0.5× published and under active-learning repair" on 2026-09-18; the repair was a sampler that ran at twice its stated temperature and a table that compared a heat-capacity peak with a free-energy crossing. Note what the fourth column does not say: v5 puts its strongest ordering on Cr–Ta in both Cr-bearing systems, where SOB20 report Cr–V. On Cr–Ta–V–W the temperature still agrees to 2 % with the wrong pair driving it, which is a warning about reading agreement as validation — and precisely why E222 tests the pair and not the temperature.

EXPERIMENTS.md · lines 15388–15391

E210b, fifth system: Cr₂₅Ti₂₅V₂₅W₂₅ — largest peak 796 / 947 K (cooling/heating), highest-T 796 / 1151 (the heating leg carries a small hot bump), floor α₁ −0.34 (Cr–Ti), E −54 meV/atom, reversible. Not scoreable yet: the table holds only its F_mix value (700 K); the H_mix ODTT is being retrieved from SOB20's Table 2. Note the strongest pair here is Cr–Ti, not Cr–Ta or Cr–V.

EXPERIMENTS.md · lines 15651–15668

E210b, systems 6–8. Cr₂₅Ta₂₅Ti₂₅V₂₅: highest-T 646 K (563/730) vs SOB20 700 → 0.92, pairs Cr–Ta −0.64 and Cr–V −0.55 (SOB20: Cr–V). Mo₂₀Nb₂₀Ta₂₀V₂₀W₂₀: 841 K (916/765; largest 733) vs FC17's 750 H_mix inflection → 1.12 (0.98 by the largest peak), pairs Mo–Ta −1.49, Mo–Nb −1.00. Mo₃₃Nb₃₃Ta₃₃: 1100 K, no resolved published value (WAK23), pairs Mo–Ta −0.87. All three reversible (junction ≤ 0.02 meV). Scorecard so far: 1.08, 1.21, 0.98, 1.22, 0.92, 1.12 — six of six comparable systems within 22 %, plus the Cr–Ta–Ti–W 2.92 that E222 explains. The ninth (MoNbTaW, the only cv_peak/chi_peak row) is running.

Seed variance, two of three seeds. v5 refitted with seed 1: held-out MAE 7.4 against seed 0's 7.2 (RMSE 10.7 vs 9.2, bias +2.6 vs +0.5). Preliminary spread ≈ 0.2 meV in MAE, so v7's 9.3 and v7b's 9.9 sit well outside it and their degradation is real, not fit noise — the pre-registered branch "a VASP-RHEA vs QE frame offset must be fitted per group before rows are merged" is the one that applies, pending seed 2 (paused by the guardian at load 13.6). E221 (the kinetics anchor) is on its second cell. The fcc reference redo landed: Fe fcc (non-spin) now fits inside its low bracket at a₀ = 3.4617 Å, E_min −4479.7572 eV/atom (the first pass had read the default 0.96–1.02 bracket, whose minimum sat below 0.96; the _lo_e80 scan at 0.93–0.99 brackets it). All five fcc references (Al, Co, Cu, Fe, Ni) are bracketed; E226 can read them.

EXPERIMENTS.md · lines 16337–16351

E210b, ninth system: MoNbTaW (Mo₂₅Nb₂₅Ta₂₅W₂₅) — v5 orders it at under half the published temperature. runs/e210b_scorecard/Mo25Nb25Ta25W25/mc.json: dE/dT peaks 247 K cooling / 321 K heating (one grid step, 73.5 K, apart); heat-capacity (variance) peaks 174 / 321 K; floors agree to 0.0004 meV/atom, so it is reversible by (1). Strongest pair Mo–Ta, α₁ at the cooling peak −1.38 (B2 on (Mo,W | Nb,Ta)). This is the scorecard's one heat-capacity-peak row and the only like-for-like comparison with the KOS19 / KW23 CE+MC work on this chemistry: 0.41× KOS19's ideal-lattice 600 K by the same observable (variance-peak mean 247 K), 0.47× by the ladder's highest-T convention (284 K), and 0.22–0.26× KW23's 1110 K susceptibility peak. Prediction (2) (0.5–1.0×) fails for MoNbTaW. The scorecard therefore reads: six of seven comparable systems within 22 % (1.08, 1.21, 0.98, 1.22, 0.92, 1.12), MoNbTaW at 0.41, and Cr–Ta–Ti–W at 2.92 (E222). The two rows that fall short are the two from CE + Monte Carlo work on Mo–Nb–Ta–W (KOS19/KW23 here; Mo–Ta at 0.49× v5 and 0.57× e6 in E242 (4)). v5 under-orders the Mo-group B2 systems while matching the SOB20 rows. The survivor is 52 % Mo. This is why E242's DFT, not rung 1, decides it. (KOS19's relaxed-lattice 250–350 K, a different ground state, is not the comparable row and is not scored.)

Related entries

Built with PRISMWebsite and visualizations made using Claude