Does the reversible ordering test give sensible temperatures on all nine published systems?
Withdrawn. It stopped after its first system; an independent second sampler ran all nine systems instead.
In the log: the reversible protocol on every scorecard system (2026-09-20 23:39; queued behind E207)
supersededDate 2026-09-20 23:39, as written in the logrung 1 · ordering3 predictions · 0 result paragraphsEXPERIMENTS.md lines 13228–13248, lines 14286–14305, lines 14628–14631
What E210 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E210.svg).
Pre-registration
(1)
and
no verdict written against it
(2)
the Cv-peak
temperatures — stand, but will be read from dE/dT as well as the variance channel, with
the grid step (2600−100)/34 = 73.5 K as the stated resolution. And four of its nine systems
(the Cr-bearing SOB20 rows) are not scoreable against the scorecard at all until SOB20's
spin treatment is known; they will be run and recorded, not scored.
A bug caught by its own test, the same hour it was written (2026-09-21 17:2x).specific_heat_two_ways had a stray /KB in the variance channel, and hysteresis()
reconstructed the variance with a matching *KB, so the E214 read-out (ratio 1.94, both
peaks 526 K) was right while the function alone was 11,604× off. A synthetic test that
builds an (E, V) pair for which the identity Var(E)/k_BT² ≡ dU/dT holds exactly returned a
ratio of 11,604 instead of 1 and would not pass. Fixed; the E214 numbers are unchanged, as
a cancelled pair of errors should leave them. Recorded because it is the third time today
that a check built on a known answer caught what a check built on agreement between two
unknowns could not.
no verdict written against it
(3)
"v5's SRO onset lands within 200 K of the H_mix-inflection rows" — was written with
the SRO threshold in pyeCE's units (half of Warren–Cowley) and with "onset" meaning a
threshold crossing on a smooth curve, which referee 2 showed is not an observable. It is
withdrawn as a scored prediction; E210's SRO curves will be reported in Warren–Cowley
units and compared as curves, not onsets. Its predictions
no verdict written against it
The pre-registration, as written
E210's pre-registration, annotated before it runs (2026-09-21 17:1x). Its prediction
(3) — "v5's SRO onset lands within 200 K of the H_mix-inflection rows" — was written with
the SRO threshold in pyeCE's units (half of Warren–Cowley) and with "onset" meaning a
threshold crossing on a smooth curve, which referee 2 showed is not an observable. It is
withdrawn as a scored prediction; E210's SRO curves will be reported in Warren–Cowley
units and compared as curves, not onsets. Its predictions (1) and (2) — the Cv-peak
temperatures — stand, but will be read from dE/dT as well as the variance channel, with
the grid step (2600−100)/34 = 73.5 K as the stated resolution. And four of its nine systems
(the Cr-bearing SOB20 rows) are not scoreable against the scorecard at all until SOB20's
spin treatment is known; they will be run and recorded, not scored.
A bug caught by its own test, the same hour it was written (2026-09-21 17:2x).specific_heat_two_ways had a stray /KB in the variance channel, and hysteresis()
reconstructed the variance with a matching *KB, so the E214 read-out (ratio 1.94, both
peaks 526 K) was right while the function alone was 11,604× off. A synthetic test that
builds an (E, V) pair for which the identity Var(E)/k_BT² ≡ dU/dT holds exactly returned a
ratio of 11,604 instead of 1 and would not pass. Fixed; the E214 numbers are unchanged, as
a cancelled pair of errors should leave them. Recorded because it is the third time today
that a check built on a known answer caught what a check built on agreement between two
unknowns could not.
E210 and every downstream number wait for the factor. E210 (nine systems) is released by
this result and runs, because a constant sampler factor leaves its relative ordering intact
and the scan is needed either way; its absolute temperatures are provisional until the
factor is pinned.
Results
No result paragraph for this entry was found in the log.
The full record
This entry is written in 3 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 13228–13248
E210 — the reversible protocol on every scorecard system (2026-09-20 23:39; queued behind E207)
Operator: "we're not just going to look at molybdenum and tantalum all our lives." Every
rung-1 test so far was Mo–Ta or MoNbTaVW because that is where the published numbers were.
E210 runs E207's two-leg protocol (6³, 25 + 25 temperatures 1400 ⇄ 100 K, 1000–2000 sweeps
per temperature, cooling then heating in one process) on all nine distinct scorecard
systems (the table's fourteen rows collapse to nine compositions) — Cr–Ta–Ti–V–W, Ta–Ti–V–W, Cr–Ta–Ti–W, Cr–Ta–V–W, Cr–Ti–V–W, Cr–Ta–Ti–V, MoNbTaVW,
Mo–Nb–Ta, MoNbTaW and Mo–Ta among them — one at a time (~3 h each, ~1.5 days). For each,
the heating-leg Cv peak and the SRO onset (|α| > 0.05) are reported with the CV-derived
error bar and scored only against like rows. Predictions: (1) every system is
reversible (floor mismatch < 3 meV, legs coincide) — if any is not, that system's number is
withheld, not reported; (2) v5's Cv-peak T_c is 0.3–0.6× the like-for-like
ideal-lattice value wherever one exists (MoNbTaW: 264 vs 600), i.e. the under-ordering is
systematic, not a MoNbTaW quirk; (3) v5's SRO onset lands within 200 K of the
H_mix-inflection rows (SOB20's "H_mix+SRO", FC17), because that is the observable those
rows actually are — which would say the old scorecard was "wrong" by comparing the wrong
pair of numbers, not because the model's onset is wrong; (4) the V-bearing systems
(B32-like per Woodgate & Staunton) show a lower v5/literature ratio than the
valence-difference B2 systems, because the model's deep-end under-fit is worst where the
ordering is driven by size rather than valence. runs/e210_chain.sh; results in
runs/e210_scorecard/<system>/hysteresis.json.
EXPERIMENTS.md · lines 14286–14305
E210's pre-registration, annotated before it runs (2026-09-21 17:1x). Its prediction
(3) — "v5's SRO onset lands within 200 K of the H_mix-inflection rows" — was written with
the SRO threshold in pyeCE's units (half of Warren–Cowley) and with "onset" meaning a
threshold crossing on a smooth curve, which referee 2 showed is not an observable. It is
withdrawn as a scored prediction; E210's SRO curves will be reported in Warren–Cowley
units and compared as curves, not onsets. Its predictions (1) and (2) — the Cv-peak
temperatures — stand, but will be read from dE/dT as well as the variance channel, with
the grid step (2600−100)/34 = 73.5 K as the stated resolution. And four of its nine systems
(the Cr-bearing SOB20 rows) are not scoreable against the scorecard at all until SOB20's
spin treatment is known; they will be run and recorded, not scored.
A bug caught by its own test, the same hour it was written (2026-09-21 17:2x).specific_heat_two_ways had a stray /KB in the variance channel, and hysteresis()
reconstructed the variance with a matching *KB, so the E214 read-out (ratio 1.94, both
peaks 526 K) was right while the function alone was 11,604× off. A synthetic test that
builds an (E, V) pair for which the identity Var(E)/k_BT² ≡ dU/dT holds exactly returned a
ratio of 11,604 instead of 1 and would not pass. Fixed; the E214 numbers are unchanged, as
a cancelled pair of errors should leave them. Recorded because it is the third time today
that a check built on a known answer caught what a check built on agreement between two
unknowns could not.
EXPERIMENTS.md · lines 14628–14631
E210 and every downstream number wait for the factor. E210 (nine systems) is released by
this result and runs, because a constant sampler factor leaves its relative ordering intact
and the scan is needed either way; its absolute temperatures are provisional until the
factor is pinned.
Related entries
E207 — does the ordering sweep equilibrate? Heating vs cooling at 10–30× the sweeps (2026-09-20…
E214 — pre-registered before it runs (2026-09-21 07:4x)