Experiments · E99

Does the cheap ordering check miss orderings in alloys away from a 50/50 mix?

Yes. For W0.73 Ta0.24 it reported no ordering, while the simulation found one at 649 K, inside the window.

In the log: Checking rung 0's ordering gate against rung 1, on the fly's own finds

confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04rung 1 · ordering0 predictions · 1 result paragraphEXPERIMENTS.md lines 5804–5868
exp E99 diagram
What E99 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E99.svg).

Results

EXPERIMENTS.md · line 5833

Outcome. All three predictions hold. The gate deployed last night is unsafe off stoichiometry and is withdrawn as a clearance.

The full record

EXPERIMENTS.md · lines 5804–5868

E99 — Checking rung 0's ordering gate against rung 1, on the fly's own finds

The gate deployed in E98 is built from a two-sublattice partition: species are split into two groups, one group on the bcc corner sites and the other on the body centres. That construction represents B2 and its relatives, which live at or near equiatomic. It cannot represent an A3B ordering, because the corner sublattice is exactly half the sites and a 0.75/0.25 composition cannot be segregated onto it.

The fly's best find is W0.73 Ta0.24, and the gate returned T_order = 0 for it - no ordering tendency at all. That is either a real property of the alloy or the construction failing to reach the structure that orders. D0_3 and A15 both sit near that stoichiometry and neither is a two-sublattice split of this cell.

Predicted:

  1. The Monte Carlo finds ordering in W0.73 Ta0.24 where the gate says none. Rung 1 samples occupations directly and is not restricted to any sublattice decomposition.
  2. The gate agrees better near equiatomic. W0.56 Ta0.44 and Ta0.56 W0.44, which the gate puts at 422 and 438 K, should come back close; the ratio should degrade as the composition moves away from 50/50.
  3. The error is therefore one-signed and unsafe - the gate under-reports ordering off stoichiometry, which is the direction that lets an orderer through.

Falsified if the Monte Carlo agrees that W0.73 Ta0.24 has no ordering transition, in which case the construction is adequate and the fly's leading candidate stands.

Rung 1 is Monte Carlo over the expansion, so it costs CPU and no potential - it can run while the 54-atom DFT holds the memory.

Outcome. All three predictions hold. The gate deployed last night is unsafe off stoichiometry and is withdrawn as a clearance.

composition x_major T_gate T_od (rung 1) ratio
Ta0.56 W0.44 0.56 339 472 0.72
W0.56 Ta0.44 0.56 398 590 0.67
W0.62 Ta0.38 0.62 217 506 0.43
W0.69 Ta0.31 0.69 135 649 0.21
W0.73 Ta0.24 0.73 0 649 0.00

The ratio degrades monotonically with distance from equiatomic and the error is one-signed: the gate always under-reports. The fly's leading candidate from the E98 campaign orders at 649 K, inside the window, and the gate cleared it. The E98 ladder campaign results are withdrawn.

A hypothesis of mine that the data refuses. I expected the generator to have found the blind spot - reward hacking one level up, moving off equiatomic because that is where the gate cannot see. It did not. The ratio of majority fraction to equiatomic is 1.50, 1.54 and 1.55 across the pessimistic, effective and ladder campaigns: unchanged whether the gate was present or not. Off-stoichiometry is simply where these alloys sit, so the gate is broken across most of the space rather than being gamed in a corner of it. That is worse as an instrument and better as a dynamic, and it was worth checking rather than asserting.

What the gate is still good for. A nonzero T_order is a real signal - the construction found an ordered state lower than random, and that state exists. A zero is not a clearance. It is kept on those terms, with the validity domain written into the module rather than left in this file.

The sub-rung, measured. Rung 1 at its full setting is 19.5 s; at ten sweeps, one seed and four temperatures it returns 687 K against the reference 649, an error of +38 K, in 1.2 seconds. That is 400 times the gate and 16 times cheaper than rung 1, which is the shape of a sub-rung: the cheap gate flags what it can see, the short Monte Carlo clears what it cannot, and the full rung 1 settles anything that matters. The open question is what the generator should learn from, since a reward it can only evaluate at 1.2 s per proposal is a different search from one at 3 ms.

Related entries

Built with PRISMWebsite and visualizations made using Claude