No result paragraph for this entry was found in the log.
EXPERIMENTS.md · lines 3211–3280E60 — The generators, compared on a real target; and a composition better than the one everyone cites
With the free rung (E59) a generator can be run for thousands of proposals, so the arms can
be compared with error bars on the question that actually decides. Target: a driving force
below -40 meV/atom at 90 K. Sixty rounds of four proposals, eight seeds, 240 verifications
per run. Every arm answers the same propose/observe contract, so a fly, a
cross-entropy method and a uniform draw are interchangeable.
Uniform sampling finds nothing at all in 240 verifications, which is what makes the
comparison worth making: the target is real and luck does not reach it.
What the fly is and is not. It beats uniform outright and beats the sparse-Dirichlet
control - its own geometry with no circuit - by 3.9 times. So the mushroom body is doing
something a corner-seeking prior does not. But it does not beat its frozen twin by any
amount that survives a test: paired over the shared seeds the difference is +0.45 +/-
0.51, t = 0.89, p = 0.40, winning on six seeds of eight. The falsifier F1 is written as a
ratio above one and 1.22 clears it, which is exactly the trap a threshold sets. Learning
is not demonstrated. And two much simpler generators beat the fly: elitist by 1.8 times,
cross-entropy by 1.3.
On the surrogate (islands placed by hand) the fly scored zero against centre targets and
matched its frozen twin on near-binary ones. On the real target it is a competent generator
that is not the best one, and the plasticity remains unproven in both settings.
A note on the falsifier arithmetic. F2 was written fly/uniform > 1.5 and reported
FAIL when uniform scored zero, because a zero denominator fell through to the failure
branch. Infinite improvement over the null is the strongest form of passing.
What was found. The elitist arm's best compositions, carried to the dearer rung - the
real MACE relaxation rather than the misfit estimate - with equiatomic MoNbTaW run under
the identical rung for comparison:
All four survive promotion, and all four are more stable against every off-lattice
competitor than the equiatomic alloy the literature cites - the best by 36 meV/atom.
Repeated under three independent sets of random occupancies, the two ends of that
comparison are -95.2 +/- 2.2 and -57.2 +/- 0.9 meV/atom, so the advantage is
38 +/- 2.4 meV/atom - about sixteen standard errors, and an error bar measured rather
than assumed.
They are consistently tantalum-rich and tungsten-poor, with (Nb+Ta) near 0.65 against
(Mo+W) near 0.35, which is not the half-and-half that would maximise B2 ordering between
those two groups. There is no explanation for that here.
The screen's error is one-signed in this corner: -6 to -22 meV/atom, mean -14, every
candidate more stable than predicted. Over the nine audited compositions it was +2 +/- 16
with no bias, so this is a local systematic rather than noise, and it runs in the
conservative direction - the screen under-promises.
What is not yet done. These are thermodynamic verdicts. None of the four has been
through the kinetic rung, and on the evidence of MoNbTaW that rung is what decides whether
a 631 K ordering matters. Tantalum-rich compositions have lower melting points than
equiatomic MoNbTaW, so their vacancies move more easily, and the kinetic margin should be
expected to be worse rather than better.