Experiments · E60

On the real target, does the fly-brain generator beat simpler generators?

No. A simple hill-climber scored 1.8 times higher, and the fly did not beat its own frozen copy (p = 0.40).

In the log: The generators, compared on a real target; and a composition better than the one everyone cites

recordedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04unclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 3211–3280
exp E60 diagram
What E60 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E60.svg).

Results

No result paragraph for this entry was found in the log.

The full record

EXPERIMENTS.md · lines 3211–3280

E60 — The generators, compared on a real target; and a composition better than the one everyone cites

With the free rung (E59) a generator can be run for thousands of proposals, so the arms can be compared with error bars on the question that actually decides. Target: a driving force below -40 meV/atom at 90 K. Sixty rounds of four proposals, eight seeds, 240 verifications per run. Every arm answers the same propose/observe contract, so a fly, a cross-entropy method and a uniform draw are interchangeable.

arm AUC_Q distinct found best (meV/atom)
uniform 0.00 +/- 0.00 0.00 -5
sparse (Dirichlet a=0.3) 0.65 +/- 0.22 1.00 -45
fly-frozen 2.09 +/- 0.28 4.38 -65
fly 2.54 +/- 0.27 6.25 -66
cem 3.24 +/- 0.82 13.00 -55
elitist 4.63 +/- 1.10 17.00 -62

Uniform sampling finds nothing at all in 240 verifications, which is what makes the comparison worth making: the target is real and luck does not reach it.

What the fly is and is not. It beats uniform outright and beats the sparse-Dirichlet control - its own geometry with no circuit - by 3.9 times. So the mushroom body is doing something a corner-seeking prior does not. But it does not beat its frozen twin by any amount that survives a test: paired over the shared seeds the difference is +0.45 +/- 0.51, t = 0.89, p = 0.40, winning on six seeds of eight. The falsifier F1 is written as a ratio above one and 1.22 clears it, which is exactly the trap a threshold sets. Learning is not demonstrated. And two much simpler generators beat the fly: elitist by 1.8 times, cross-entropy by 1.3.

On the surrogate (islands placed by hand) the fly scored zero against centre targets and matched its frozen twin on near-binary ones. On the real target it is a competent generator that is not the best one, and the plasticity remains unproven in both settings.

A note on the falsifier arithmetic. F2 was written fly/uniform > 1.5 and reported FAIL when uniform scored zero, because a zero denominator fell through to the failure branch. Infinite improvement over the null is the strongest form of passing.

What was found. The elitist arm's best compositions, carried to the dearer rung - the real MACE relaxation rather than the misfit estimate - with equiatomic MoNbTaW run under the identical rung for comparison:

composition screen MACE relaxed
Ta.40 Mo.26 Nb.25 W.09 -70 -93
Mo.37 Ta.34 Nb.24 W.05 -68 -83
Mo.41 Ta.30 Nb.29 -69 -81
Ta.44 W.27 Mo.21 Nb.08 -64 -79
Mo.25 Nb.25 Ta.25 W.25 -51 -57

All four survive promotion, and all four are more stable against every off-lattice competitor than the equiatomic alloy the literature cites - the best by 36 meV/atom.

Repeated under three independent sets of random occupancies, the two ends of that comparison are -95.2 +/- 2.2 and -57.2 +/- 0.9 meV/atom, so the advantage is 38 +/- 2.4 meV/atom - about sixteen standard errors, and an error bar measured rather than assumed. They are consistently tantalum-rich and tungsten-poor, with (Nb+Ta) near 0.65 against (Mo+W) near 0.35, which is not the half-and-half that would maximise B2 ordering between those two groups. There is no explanation for that here.

The screen's error is one-signed in this corner: -6 to -22 meV/atom, mean -14, every candidate more stable than predicted. Over the nine audited compositions it was +2 +/- 16 with no bias, so this is a local systematic rather than noise, and it runs in the conservative direction - the screen under-promises.

What is not yet done. These are thermodynamic verdicts. None of the four has been through the kinetic rung, and on the evidence of MoNbTaW that rung is what decides whether a 631 K ordering matters. Tantalum-rich compositions have lower melting points than equiatomic MoNbTaW, so their vacancies move more easily, and the kinetic margin should be expected to be worse rather than better.

Related entries

Built with PRISMWebsite and visualizations made using Claude