Experiments · E71

Does the fly's learning advantage survive when its reward is corrected?

Yes. It still beat its frozen copy (p = 0.0006), but the advantage shrank from 3.9 to 1.7 times.

In the log: The learning result survives its reward being replaced

confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04generator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 3880–3921
exp E71 diagram
What E71 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E71.svg).

Results

EXPERIMENTS.md · line 3880

E71 — The learning result survives its reward being replaced

E70 rebuilt the relaxation model, and that model feeds the screen, and the screen is the reward the generator learns from. Every result in E60 and E61 was measured against the version now known to have had no skill. The load-bearing one is E61: the fly beating its own frozen twin on sixteen seeds of sixteen.

Predicted before rerunning: it survives, because the mechanism does not depend on the particular reward, only on there being a learnable signal, and the corrected screen still ranks the same corner highest. A failure would mean E61 was an artefact of a broken reward.

old reward corrected reward
uniform 0.00 0.09 +/- 0.04
sparse 0.65 4.78 +/- 0.52
cem 3.24 13.28 +/- 2.71
elitist 4.63 16.13 +/- 2.87
fly-frozen 1.61 11.10 +/- 0.61
fly 6.29 19.06 +/- 1.16

Paired fly minus frozen +7.96 +/- 1.52, t one-tailed p = 0.0006, Wilcoxon p = 0.008, seven of eight seeds. More significant than before, and the result stands.

But the ratio shrank, from 3.92 to 1.72, and the reason matters. The frozen twin went from 1.61 to 11.10. The corrected screen makes the target five times less rare (E59, as annotated), so run-and-tumble on a corner-seeking geometry now finds a great deal on its own and the learning has proportionally less left to add. The effect is more certain and smaller, which is what happens when a problem gets easier.

Two claims of this project's own are corrected by the same run.

"Uniform sampling finds nothing." Repeated several times, and it was partly an artefact: the -40 meV/atom target was one that no uniform draw in twenty thousand could reach under the broken screen. Corrected, uniform finds 0.38 distinct compositions against the elitist hill-climber's 50. Still hopeless, and not zero.

The fly against the hill-climber. 19.06 against 16.13, paired p = 0.172. Still not significant, as in E61. Three separate rewards now, and the fly has never been shown to beat a hill-climber on this problem - only to beat its own frozen twin, which is the claim about learning rather than about superiority.

One seed of eight went the other way this time, which did not happen under the old reward.


The full record

EXPERIMENTS.md · lines 3880–3921

E71 — The learning result survives its reward being replaced

E70 rebuilt the relaxation model, and that model feeds the screen, and the screen is the reward the generator learns from. Every result in E60 and E61 was measured against the version now known to have had no skill. The load-bearing one is E61: the fly beating its own frozen twin on sixteen seeds of sixteen.

Predicted before rerunning: it survives, because the mechanism does not depend on the particular reward, only on there being a learnable signal, and the corrected screen still ranks the same corner highest. A failure would mean E61 was an artefact of a broken reward.

old reward corrected reward
uniform 0.00 0.09 +/- 0.04
sparse 0.65 4.78 +/- 0.52
cem 3.24 13.28 +/- 2.71
elitist 4.63 16.13 +/- 2.87
fly-frozen 1.61 11.10 +/- 0.61
fly 6.29 19.06 +/- 1.16

Paired fly minus frozen +7.96 +/- 1.52, t one-tailed p = 0.0006, Wilcoxon p = 0.008, seven of eight seeds. More significant than before, and the result stands.

But the ratio shrank, from 3.92 to 1.72, and the reason matters. The frozen twin went from 1.61 to 11.10. The corrected screen makes the target five times less rare (E59, as annotated), so run-and-tumble on a corner-seeking geometry now finds a great deal on its own and the learning has proportionally less left to add. The effect is more certain and smaller, which is what happens when a problem gets easier.

Two claims of this project's own are corrected by the same run.

"Uniform sampling finds nothing." Repeated several times, and it was partly an artefact: the -40 meV/atom target was one that no uniform draw in twenty thousand could reach under the broken screen. Corrected, uniform finds 0.38 distinct compositions against the elitist hill-climber's 50. Still hopeless, and not zero.

The fly against the hill-climber. 19.06 against 16.13, paired p = 0.172. Still not significant, as in E61. Three separate rewards now, and the fly has never been shown to beat a hill-climber on this problem - only to beat its own frozen twin, which is the claim about learning rather than about superiority.

One seed of eight went the other way this time, which did not happen under the old reward.

Related entries

Built with PRISMWebsite and visualizations made using Claude