With foreign elements removed, does the fly brain beat simple searches that use no brain?
Withdrawn. It tied keep-the-best-and-mutate on finds (158 against 169.5) and lost on rate, but this reversed once the reward's gate was removed.
In the log: first arm (2026-09-19 00:06)
supersededDate 2026-09-19 00:06, as written in the loggenerator · fly brain0 predictions · 4 result paragraphsEXPERIMENTS.md line 12570, line 12584, line 12586, lines 12588–12609
What E195b did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E195b.svg).
Results
EXPERIMENTS.md · line 12570
E195b, first arm (2026-09-19 00:06). uniform on the eight-element space: AUC_Q 3.37 ± 1.18, 8.5 ± 1.5 finds, best −74 — the refusal confound is gone (12-element: 0 finds). The bar every learning arm must clear is now a real one.
EXPERIMENTS.md · line 12584
E195b, third arm (2026-09-19 00:18). elitist: AUC_Q 88.90 ± 0.39, 169.5 ± 10.5 finds, best −104. Model-free order so far 8.5 < 34.5 < 169.5. The fly's best on any space to date is 77 finds (E193, one seed); prediction 1 needs it above ~180 here.
EXPERIMENTS.md · line 12586
E195b, fourth arm (2026-09-19 00:27). cem: AUC_Q 76.07 ± 2.11, 128 ± 0 finds, best −101. Order so far 8.5 < 34.5 < 169.5 > 128: elitist beats CEM again (prediction 2's last inequality fails on both spaces). Fly running; prediction 1 bar = 128 + one seed-sd.
EXPERIMENTS.md · line 12588
E195b result (2026-09-19 01:12) — the clean five-arm comparison, eight-element space, v5 reward.
arm
AUC_Q
distinct finds
best v5
mean v5 of finds
uniform
3.37 ± 1.18
8.5 ± 1.5
−74
—
sparse
14.48 ± 0.42
34.5 ± 3.5
−100
—
elitist
88.90 ± 0.39
169.5 ± 10.5
−104
—
cem
76.07 ± 2.11
128 ± 0
−105.5
−73.1
fly
73.54 ± 16.25
158 ± 17
−103
−61.9
P1 confirmed at the margin (fly − cem = +30 finds vs one seed-sd 24, ddof 1): 1.25 sd,
two seeds. P2 falsified (elitist > cem on both spaces). P3 falsified (fly's finds
11 meV shallower than CEM's). P4 ✓ (five arms in 75 min). The headline: the
whole-brain head with a live gain field ties elitist on finds (158 ± 17 vs 169.5 ± 10.5)
and loses to both elitist and CEM on AUC_Q — the rate of distinct finds per unit spend,
which is the metric the generator line was built on (trial.py) — with the widest seed
spread of any arm and the shallowest finds. On a reward that ranks, the fly's advantage
over model-free search (E136, E185: flat reward and island surrogate) does not appear.
What survives: the fly beats uniform and sparse by an order of magnitude and CEM on count.
What does not: "more distinct qualifiers per unit of verifier cost than the baselines."
Third seeds for fly and elitist (E195c) are queued to settle the count tie; the AUC gap
(73.5 vs 88.9, four seed-sd of elitist) will not close with a third seed.
The full record
This entry is written in 4 separate places in the log, shown here in log order.
EXPERIMENTS.md · line 12570
E195b, first arm (2026-09-19 00:06). uniform on the eight-element space: AUC_Q 3.37 ± 1.18, 8.5 ± 1.5 finds, best −74 — the refusal confound is gone (12-element: 0 finds). The bar every learning arm must clear is now a real one.
EXPERIMENTS.md · line 12584
E195b, third arm (2026-09-19 00:18). elitist: AUC_Q 88.90 ± 0.39, 169.5 ± 10.5 finds, best −104. Model-free order so far 8.5 < 34.5 < 169.5. The fly's best on any space to date is 77 finds (E193, one seed); prediction 1 needs it above ~180 here.
EXPERIMENTS.md · line 12586
E195b, fourth arm (2026-09-19 00:27). cem: AUC_Q 76.07 ± 2.11, 128 ± 0 finds, best −101. Order so far 8.5 < 34.5 < 169.5 > 128: elitist beats CEM again (prediction 2's last inequality fails on both spaces). Fly running; prediction 1 bar = 128 + one seed-sd.
EXPERIMENTS.md · lines 12588–12609
E195b result (2026-09-19 01:12) — the clean five-arm comparison, eight-element space, v5 reward.
arm
AUC_Q
distinct finds
best v5
mean v5 of finds
uniform
3.37 ± 1.18
8.5 ± 1.5
−74
—
sparse
14.48 ± 0.42
34.5 ± 3.5
−100
—
elitist
88.90 ± 0.39
169.5 ± 10.5
−104
—
cem
76.07 ± 2.11
128 ± 0
−105.5
−73.1
fly
73.54 ± 16.25
158 ± 17
−103
−61.9
P1 confirmed at the margin (fly − cem = +30 finds vs one seed-sd 24, ddof 1): 1.25 sd,
two seeds. P2 falsified (elitist > cem on both spaces). P3 falsified (fly's finds
11 meV shallower than CEM's). P4 ✓ (five arms in 75 min). The headline: the
whole-brain head with a live gain field ties elitist on finds (158 ± 17 vs 169.5 ± 10.5)
and loses to both elitist and CEM on AUC_Q — the rate of distinct finds per unit spend,
which is the metric the generator line was built on (trial.py) — with the widest seed
spread of any arm and the shallowest finds. On a reward that ranks, the fly's advantage
over model-free search (E136, E185: flat reward and island surrogate) does not appear.
What survives: the fly beats uniform and sparse by an order of magnitude and CEM on count.
What does not: "more distinct qualifiers per unit of verifier cost than the baselines."
Third seeds for fly and elitist (E195c) are queued to settle the count tie; the AUC gap
(73.5 vs 88.9, four seed-sd of elitist) will not close with a third seed.
Related entries
E193 — the search on the non-flat reward (2026-09-18 16:29; operator: icet retired from the…
E136 — The fly has not been used, and nothing found so far is new
E185 — the heterogeneity question as one MPS family (E167 / E170 / E174, same device)