Does rewarding low energy steer the search toward alloys that fail the ordering test?
No. The opposite: top-rewarded compositions pass both tests 12 times more often than the grid overall; later corrections raised this to 34.
In the log: Is rung 0's reward anti-aligned with the actual requirement?
falsifiedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04rung 1 · ordering0 predictions · 1 result paragraphEXPERIMENTS.md lines 6898–6937, lines 6939–6989
What E117 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E117.svg).
Results
EXPERIMENTS.md · line 6939
E117 result: prediction 1 confirmed, predictions 2 and 3 FALSIFIED. The objective is not
anti-aligned - it is a 12-fold enrichment. Fable's central claim does not survive
measurement.
1705 grid compositions, one rung-0 screen each.
Prediction 1 confirmed.Mo0.50 Ta0.50 at -86.6 meV/atom is the global minimum of
drive_conservative over the entire grid. The search was never broken. It converges on
the objective's true optimum, and E114's and E116's collapse diagnosis - correct as
description - blamed the wrong component.
Predictions 2 and 3 falsified, and the first pass of this analysis was wrong too.p_survives_window is 1 - p_happens, and p_happens is the probability of decomposition
into off-lattice competitors, not of the order-disorder transition. Reading it as the
phase requirement conflated two independent failure modes. The requirement needs both:
no decomposition (p_survives_window > 0.9) and no ordering transition inside the window
(orders_in_window false).
Scored properly:
passes BOTH tests, whole grid 55 of 1705 = 3.2%
passes BOTH, best 50 by reward = 38%
Spearman(drive_conservative, passes) -0.250
A twelve-fold enrichment. Rewarding driving force is a good filter for the actual
requirement, not a bad one. Fable's claim that "mixing enthalpy and ordering tendency are the
same number, so the objective steers towards what rung 1 rejects" is falsified for this
grid: the correlation runs the other way.
What is genuinely wrong is the reporting, not the search or the objective. Fifty-five
compositions pass, and the generator reports the top few. The diversity that was missing is
an archive over the feasible set, which is the weaker "multi-basin reporting" use of
MAP-Elites, not MAP-Elites as a replacement optimiser. That is consistent with E114, where
MAP-Elites-as-optimiser lost by 98.8 meV/atom.
Widening to twelve elements paid off, measurably. Eleven of the 55 passing compositions
contain Cr, Ni, Cu or Co, including Cr0.75 Ni0.25 at -30.2 meV/atom with no transition in
the window. The eight-element space could not represent any of them.
Unverified claims from the same review, recorded because they would outrank everything
above if true. None is in any rung, and none has been checked here:
MoO3 volatilises above ~800 K, so a Mo-rich alloy may not survive 1000 K in oxygen at all.
Ti, Zr and Hf are ruled out for liquid-oxygen service by impact ignition (NASA-STD-6001).
Mo and W are brittle at 90 K (ductile-brittle transition ~300-450 K and ~500-700 K), so a
Mo-Ta part that never changes phase can still crack on thermal cycling.
If those hold, the requirement is not a phase-stability problem at all for this element set,
and Mo-Ta - which this experiment confirms is the correct answer to the question as posed -
is the wrong answer to the question that matters. Checking them is the next task, and it is
cheap: they are literature constants, not simulations.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 6898–6937
E117 — Is rung 0's reward anti-aligned with the actual requirement?
Fable's review makes a claim that, if true, invalidates the whole generator loop rather than
its search: mixing enthalpy and ordering tendency are the same number. Strong unlike-pair
attraction is what puts a composition below the hull AND what drives B2 ordering. The
generator maximises -drive_conservative, so it is steered towards the strongest orderer -
which is precisely what rung 1 then rejects for ordering inside the 90-1000 K window.
If that holds, the ladder's first rung is selecting against its own second rung's criterion,
and no amount of exploration fixes it. It is cheap to test: screen() already returns both
drive_conservative and T_order for the same composition.
The related bookkeeping claim is already confirmed: the off-lattice hull holds 5 Mo-Ta
binaries and all five are Laves (C14/C15) or Mo2Ta - no bcc-ordered Mo-Ta at all, because
E85's coordination-number filter removed the B2. That filter is defensible on its own terms,
since the expansion is supposed to cover the bcc lattice and the hull the off-lattice
competitors, so the two must not double-count. But it means drive_conservative answers
"does this beat the Laves phases", not "is this the ground state", and the ordering question
is delegated entirely to rung 1.
Enumerating the composition grid settles Fable's other question at the same time - whether the
search is broken or the objective is: all 28 binaries at 25/50/75, all equiatomic ternaries,
quaternaries and quinaries of the 12 elements, one rung-0 screen each.
Predicted:
Mo-Ta near 50:50 is at or very near the minimum of drive_conservative over the whole
grid. If so the search was never broken: it found the objective's true optimum, and
E114's and E116's collapse diagnosis, while correct as description, blamed the wrong
component.
drive_conservative and T_order are strongly NEGATIVELY correlated - the more
negative the driving force, the higher the ordering temperature. Spearman rho below -0.5.
Consequently the top-ranked compositions by reward are disproportionately those that
order INSIDE the 90-1000 K window, i.e. the ones that fail the requirement. Measured as:
the fraction of the best-50 by reward with orders_in_window true exceeds the fraction
over the whole grid.
Falsified if rho is near zero or positive, in which case the driving force and the
ordering temperature are independent, rung 0 is a legitimate cheap filter, and the problem
really is the search after all.
EXPERIMENTS.md · lines 6939–6989
E117 result: prediction 1 confirmed, predictions 2 and 3 FALSIFIED. The objective is not
anti-aligned - it is a 12-fold enrichment. Fable's central claim does not survive
measurement.
1705 grid compositions, one rung-0 screen each.
Prediction 1 confirmed.Mo0.50 Ta0.50 at -86.6 meV/atom is the global minimum of
drive_conservative over the entire grid. The search was never broken. It converges on
the objective's true optimum, and E114's and E116's collapse diagnosis - correct as
description - blamed the wrong component.
Predictions 2 and 3 falsified, and the first pass of this analysis was wrong too.p_survives_window is 1 - p_happens, and p_happens is the probability of decomposition
into off-lattice competitors, not of the order-disorder transition. Reading it as the
phase requirement conflated two independent failure modes. The requirement needs both:
no decomposition (p_survives_window > 0.9) and no ordering transition inside the window
(orders_in_window false).
Scored properly:
passes BOTH tests, whole grid 55 of 1705 = 3.2%
passes BOTH, best 50 by reward = 38%
Spearman(drive_conservative, passes) -0.250
A twelve-fold enrichment. Rewarding driving force is a good filter for the actual
requirement, not a bad one. Fable's claim that "mixing enthalpy and ordering tendency are the
same number, so the objective steers towards what rung 1 rejects" is falsified for this
grid: the correlation runs the other way.
What is genuinely wrong is the reporting, not the search or the objective. Fifty-five
compositions pass, and the generator reports the top few. The diversity that was missing is
an archive over the feasible set, which is the weaker "multi-basin reporting" use of
MAP-Elites, not MAP-Elites as a replacement optimiser. That is consistent with E114, where
MAP-Elites-as-optimiser lost by 98.8 meV/atom.
Widening to twelve elements paid off, measurably. Eleven of the 55 passing compositions
contain Cr, Ni, Cu or Co, including Cr0.75 Ni0.25 at -30.2 meV/atom with no transition in
the window. The eight-element space could not represent any of them.
Unverified claims from the same review, recorded because they would outrank everything
above if true. None is in any rung, and none has been checked here:
MoO3 volatilises above ~800 K, so a Mo-rich alloy may not survive 1000 K in oxygen at all.
Ti, Zr and Hf are ruled out for liquid-oxygen service by impact ignition (NASA-STD-6001).
Mo and W are brittle at 90 K (ductile-brittle transition ~300-450 K and ~500-700 K), so a
Mo-Ta part that never changes phase can still crack on thermal cycling.
If those hold, the requirement is not a phase-stability problem at all for this element set,
and Mo-Ta - which this experiment confirms is the correct answer to the question as posed -
is the wrong answer to the question that matters. Checking them is the next task, and it is
cheap: they are literature constants, not simulations.