Experiments · E142

Is the quick ordering gauge reliable for equal-share alloys, as earlier assumed?

No. It under-read all three, at 0.09 to 0.77 of the simulated value, so no single correction factor rescues it.

In the log: E99 certified the ordering gate at equiatomic from data taken off equiatomic

confirmedDate not stated in the log; it was written between the commit of 2026-09-16 15:53 and the first commit that contains it, 2026-09-16 16:53rung 1 · ordering0 predictions · 1 result paragraphEXPERIMENTS.md lines 8633–8659, lines 8661–8713
exp E142 diagram
What E142 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E142.svg).

Results

EXPERIMENTS.md · line 8661

E142 result: all three predictions confirmed, and E99's validity domain is parameterised on the wrong variable entirely.

Three equiatomic compositions, rung-0 gate against rung-1 Monte Carlo, expressed as E99 expressed it — gate divided by truth, so 1.00 is perfect and lower is worse:

alloy        elements   x_major   gate    rung 1   gate/truth
MoNbTaW             4      0.25    475       637         0.77
MoNbTaVW            5      0.20    323       691         0.48
MoNbTaTiW           5      0.20    158      1763         0.09

E99's own measurements:   0.72 at x_major 0.56,  0.21 at 0.69,  0.00 at 0.73

Prediction 1 confirmed: the gate under-reports at equiatomic, every one. Prediction 2 confirmed: the ratios span 0.09 to 0.77, a factor of 8.5. No single scaling rescues the gate. Prediction 3 confirmed in the ordering, though MoNbTaW rather than MoNbTaVW is the mildest.

The finding is bigger than "E99 extrapolated". E99's story is that the gate degrades as composition moves away from equiatomic, so it should be best at low x_major. Put its points and these together:

x_major   0.20   0.20   0.25   0.56   0.69   0.73
ratio     0.09   0.48   0.77   0.72   0.21   0.00

At x_major 0.20 the gate is both the best and the worst it has ever measured. The relation is not monotonic in x_major and x_major is not what governs it. E99's validity domain is not merely extrapolated past its data — it is built on the wrong variable.

What the three points do line up with is element count, and titanium specifically: the four-element alloy is fine at 0.77, both five-element alloys are worse, and the one containing Ti is catastrophic at 0.09. That is consistent with ordering.py's own documented limit — the gate is a two-sublattice B2 construction and returns a lower bound, and the more distinct species there are the less of the real ordering two sublattices can hold.

Consequences.

  • E99's certification is withdrawn. Not the withdrawal of the gate as a clearance, which stands and is if anything strengthened — what is withdrawn is the claim that the gate can be trusted near equiatomic.
  • E131 and E132 kept twelve compositions on the grounds that a nonzero T_order is a real signal. That grounding is gone. Those twelve carry readings from an instrument now measured at 0.09 to 0.77 of truth with no calibration, so the two survivors of E132 — Mo0.50 Ta0.50 and Ta0.75 W0.25 — are not established either. The only compositions whose ordering is actually known are the three measured at rung 1.
  • The direction is safe. The gate under-reports, so it puts transitions lower than they are, which pushes compositions into the window and rejects them. Nothing has been passed that should have failed; things have been failed that should have passed. MoNbTaTiW is exactly that case, and there may be more.

This is the third instrument in this project whose stated validity domain did not survive measurement — after the Q proxy (E92) and the flat baseline (E110). The pattern is that a domain gets asserted from the data in hand and then used outside it.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 8633–8659

E142 — E99 certified the ordering gate at equiatomic from data taken off equiatomic

E99 withdrew the rung-0 ordering gate as a clearance and drew its validity domain from a sweep in major-element fraction: ratio 0.72 at x_major 0.56, 0.21 at 0.69, 0.00 at 0.73. It concluded the gate degrades away from equiatomic and is reliable at it. Every consumer since has relied on that, including E131 and E132, which kept the twelve compositions whose T_order was nonzero on the grounds that the reading could be trusted.

But E99 never measured an equiatomic composition. Its coldest point is x_major 0.56, and the certification at equiatomic is an extrapolation off the end of its own data.

Three equiatomic compositions now have both a rung-0 gate value and a rung-1 Monte Carlo value, which is the comparison E99 never made. This uses only data already on disk.

Predicted:

  1. The gate under-reports at equiatomic too, so E99's validity domain is wrong at the one place it mattered.
  2. The ratio is not constant — it varies by more than a factor of five across the three, so there is no calibration that would rescue the gate by scaling it.
  3. MoNbTaTiW is the worst case, at roughly 14x, and MoNbTaVW the mildest, around 2x, because the five-element alloys differ in how much of their ordering the two-sublattice construction can represent.

Falsified if the ratios cluster near one, which would mean the gate is fine at equiatomic, E99's extrapolation happened to hold, and MoNbTaTiW's 121 K was an isolated failure rather than a systematic one.

EXPERIMENTS.md · lines 8661–8713

E142 result: all three predictions confirmed, and E99's validity domain is parameterised on the wrong variable entirely.

Three equiatomic compositions, rung-0 gate against rung-1 Monte Carlo, expressed as E99 expressed it — gate divided by truth, so 1.00 is perfect and lower is worse:

alloy        elements   x_major   gate    rung 1   gate/truth
MoNbTaW             4      0.25    475       637         0.77
MoNbTaVW            5      0.20    323       691         0.48
MoNbTaTiW           5      0.20    158      1763         0.09

E99's own measurements:   0.72 at x_major 0.56,  0.21 at 0.69,  0.00 at 0.73

Prediction 1 confirmed: the gate under-reports at equiatomic, every one. Prediction 2 confirmed: the ratios span 0.09 to 0.77, a factor of 8.5. No single scaling rescues the gate. Prediction 3 confirmed in the ordering, though MoNbTaW rather than MoNbTaVW is the mildest.

The finding is bigger than "E99 extrapolated". E99's story is that the gate degrades as composition moves away from equiatomic, so it should be best at low x_major. Put its points and these together:

x_major   0.20   0.20   0.25   0.56   0.69   0.73
ratio     0.09   0.48   0.77   0.72   0.21   0.00

At x_major 0.20 the gate is both the best and the worst it has ever measured. The relation is not monotonic in x_major and x_major is not what governs it. E99's validity domain is not merely extrapolated past its data — it is built on the wrong variable.

What the three points do line up with is element count, and titanium specifically: the four-element alloy is fine at 0.77, both five-element alloys are worse, and the one containing Ti is catastrophic at 0.09. That is consistent with ordering.py's own documented limit — the gate is a two-sublattice B2 construction and returns a lower bound, and the more distinct species there are the less of the real ordering two sublattices can hold.

Consequences.

  • E99's certification is withdrawn. Not the withdrawal of the gate as a clearance, which stands and is if anything strengthened — what is withdrawn is the claim that the gate can be trusted near equiatomic.
  • E131 and E132 kept twelve compositions on the grounds that a nonzero T_order is a real signal. That grounding is gone. Those twelve carry readings from an instrument now measured at 0.09 to 0.77 of truth with no calibration, so the two survivors of E132 — Mo0.50 Ta0.50 and Ta0.75 W0.25 — are not established either. The only compositions whose ordering is actually known are the three measured at rung 1.
  • The direction is safe. The gate under-reports, so it puts transitions lower than they are, which pushes compositions into the window and rejects them. Nothing has been passed that should have failed; things have been failed that should have passed. MoNbTaTiW is exactly that case, and there may be more.

This is the third instrument in this project whose stated validity domain did not survive measurement — after the Q proxy (E92) and the flat baseline (E110). The pattern is that a domain gets asserted from the data in hand and then used outside it.

Related entries

Built with PRISMWebsite and visualizations made using Claude