Does the fly's learning advantage survive when its reward is corrected?
Yes. It still beat its frozen copy (p = 0.0006), but the advantage shrank from 3.9 to 1.7 times.
In the log: The learning result survives its reward being replaced
confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04generator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 3880–3921
What E71 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E71.svg).
Results
EXPERIMENTS.md · line 3880
E71 — The learning result survives its reward being replaced
E70 rebuilt the relaxation model, and that model feeds the screen, and the screen is the
reward the generator learns from. Every result in E60 and E61 was measured against the
version now known to have had no skill. The load-bearing one is E61: the fly beating its own
frozen twin on sixteen seeds of sixteen.
Predicted before rerunning: it survives, because the mechanism does not depend on the
particular reward, only on there being a learnable signal, and the corrected screen still
ranks the same corner highest. A failure would mean E61 was an artefact of a broken reward.
old reward
corrected reward
uniform
0.00
0.09 +/- 0.04
sparse
0.65
4.78 +/- 0.52
cem
3.24
13.28 +/- 2.71
elitist
4.63
16.13 +/- 2.87
fly-frozen
1.61
11.10 +/- 0.61
fly
6.29
19.06 +/- 1.16
Paired fly minus frozen +7.96 +/- 1.52, t one-tailed p = 0.0006, Wilcoxon p = 0.008, seven
of eight seeds. More significant than before, and the result stands.
But the ratio shrank, from 3.92 to 1.72, and the reason matters. The frozen twin went
from 1.61 to 11.10. The corrected screen makes the target five times less rare (E59, as
annotated), so run-and-tumble on a corner-seeking geometry now finds a great deal on its own
and the learning has proportionally less left to add. The effect is more certain and
smaller, which is what happens when a problem gets easier.
Two claims of this project's own are corrected by the same run.
"Uniform sampling finds nothing." Repeated several times, and it was partly an artefact:
the -40 meV/atom target was one that no uniform draw in twenty thousand could reach under
the broken screen. Corrected, uniform finds 0.38 distinct compositions against the elitist
hill-climber's 50. Still hopeless, and not zero.
The fly against the hill-climber. 19.06 against 16.13, paired p = 0.172. Still not
significant, as in E61. Three separate rewards now, and the fly has never been shown to beat
a hill-climber on this problem - only to beat its own frozen twin, which is the claim about
learning rather than about superiority.
One seed of eight went the other way this time, which did not happen under the old reward.
The full record
EXPERIMENTS.md · lines 3880–3921
E71 — The learning result survives its reward being replaced
E70 rebuilt the relaxation model, and that model feeds the screen, and the screen is the
reward the generator learns from. Every result in E60 and E61 was measured against the
version now known to have had no skill. The load-bearing one is E61: the fly beating its own
frozen twin on sixteen seeds of sixteen.
Predicted before rerunning: it survives, because the mechanism does not depend on the
particular reward, only on there being a learnable signal, and the corrected screen still
ranks the same corner highest. A failure would mean E61 was an artefact of a broken reward.
old reward
corrected reward
uniform
0.00
0.09 +/- 0.04
sparse
0.65
4.78 +/- 0.52
cem
3.24
13.28 +/- 2.71
elitist
4.63
16.13 +/- 2.87
fly-frozen
1.61
11.10 +/- 0.61
fly
6.29
19.06 +/- 1.16
Paired fly minus frozen +7.96 +/- 1.52, t one-tailed p = 0.0006, Wilcoxon p = 0.008, seven
of eight seeds. More significant than before, and the result stands.
But the ratio shrank, from 3.92 to 1.72, and the reason matters. The frozen twin went
from 1.61 to 11.10. The corrected screen makes the target five times less rare (E59, as
annotated), so run-and-tumble on a corner-seeking geometry now finds a great deal on its own
and the learning has proportionally less left to add. The effect is more certain and
smaller, which is what happens when a problem gets easier.
Two claims of this project's own are corrected by the same run.
"Uniform sampling finds nothing." Repeated several times, and it was partly an artefact:
the -40 meV/atom target was one that no uniform draw in twenty thousand could reach under
the broken screen. Corrected, uniform finds 0.38 distinct compositions against the elitist
hill-climber's 50. Still hopeless, and not zero.
The fly against the hill-climber. 19.06 against 16.13, paired p = 0.172. Still not
significant, as in E61. Three separate rewards now, and the fly has never been shown to beat
a hill-climber on this problem - only to beat its own frozen twin, which is the claim about
learning rather than about superiority.
One seed of eight went the other way this time, which did not happen under the old reward.