Experiments · E195b

With foreign elements removed, does the fly brain beat simple searches that use no brain?

Withdrawn. It tied keep-the-best-and-mutate on finds (158 against 169.5) and lost on rate, but this reversed once the reward's gate was removed.

In the log: first arm (2026-09-19 00:06)

supersededDate 2026-09-19 00:06, as written in the loggenerator · fly brain0 predictions · 4 result paragraphsEXPERIMENTS.md line 12570, line 12584, line 12586, lines 12588–12609
exp E195b diagram
What E195b did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E195b.svg).

Results

EXPERIMENTS.md · line 12570

E195b, first arm (2026-09-19 00:06). uniform on the eight-element space: AUC_Q 3.37 ± 1.18, 8.5 ± 1.5 finds, best −74 — the refusal confound is gone (12-element: 0 finds). The bar every learning arm must clear is now a real one.

EXPERIMENTS.md · line 12584

E195b, third arm (2026-09-19 00:18). elitist: AUC_Q 88.90 ± 0.39, 169.5 ± 10.5 finds, best −104. Model-free order so far 8.5 < 34.5 < 169.5. The fly's best on any space to date is 77 finds (E193, one seed); prediction 1 needs it above ~180 here.

EXPERIMENTS.md · line 12586

E195b, fourth arm (2026-09-19 00:27). cem: AUC_Q 76.07 ± 2.11, 128 ± 0 finds, best −101. Order so far 8.5 < 34.5 < 169.5 > 128: elitist beats CEM again (prediction 2's last inequality fails on both spaces). Fly running; prediction 1 bar = 128 + one seed-sd.

EXPERIMENTS.md · line 12588

E195b result (2026-09-19 01:12) — the clean five-arm comparison, eight-element space, v5 reward.

arm AUC_Q distinct finds best v5 mean v5 of finds
uniform 3.37 ± 1.18 8.5 ± 1.5 −74 —
sparse 14.48 ± 0.42 34.5 ± 3.5 −100 —
elitist 88.90 ± 0.39 169.5 ± 10.5 −104 —
cem 76.07 ± 2.11 128 ± 0 −105.5 −73.1
fly 73.54 ± 16.25 158 ± 17 −103 −61.9

P1 confirmed at the margin (fly − cem = +30 finds vs one seed-sd 24, ddof 1): 1.25 sd, two seeds. P2 falsified (elitist > cem on both spaces). P3 falsified (fly's finds 11 meV shallower than CEM's). P4 ✓ (five arms in 75 min). The headline: the whole-brain head with a live gain field ties elitist on finds (158 ± 17 vs 169.5 ± 10.5) and loses to both elitist and CEM on AUC_Q — the rate of distinct finds per unit spend, which is the metric the generator line was built on (trial.py) — with the widest seed spread of any arm and the shallowest finds. On a reward that ranks, the fly's advantage over model-free search (E136, E185: flat reward and island surrogate) does not appear. What survives: the fly beats uniform and sparse by an order of magnitude and CEM on count. What does not: "more distinct qualifiers per unit of verifier cost than the baselines." Third seeds for fly and elitist (E195c) are queued to settle the count tie; the AUC gap (73.5 vs 88.9, four seed-sd of elitist) will not close with a third seed.

The full record

This entry is written in 4 separate places in the log, shown here in log order.

EXPERIMENTS.md · line 12570

E195b, first arm (2026-09-19 00:06). uniform on the eight-element space: AUC_Q 3.37 ± 1.18, 8.5 ± 1.5 finds, best −74 — the refusal confound is gone (12-element: 0 finds). The bar every learning arm must clear is now a real one.

EXPERIMENTS.md · line 12584

E195b, third arm (2026-09-19 00:18). elitist: AUC_Q 88.90 ± 0.39, 169.5 ± 10.5 finds, best −104. Model-free order so far 8.5 < 34.5 < 169.5. The fly's best on any space to date is 77 finds (E193, one seed); prediction 1 needs it above ~180 here.

EXPERIMENTS.md · line 12586

E195b, fourth arm (2026-09-19 00:27). cem: AUC_Q 76.07 ± 2.11, 128 ± 0 finds, best −101. Order so far 8.5 < 34.5 < 169.5 > 128: elitist beats CEM again (prediction 2's last inequality fails on both spaces). Fly running; prediction 1 bar = 128 + one seed-sd.

EXPERIMENTS.md · lines 12588–12609

E195b result (2026-09-19 01:12) — the clean five-arm comparison, eight-element space, v5 reward.

arm AUC_Q distinct finds best v5 mean v5 of finds
uniform 3.37 ± 1.18 8.5 ± 1.5 −74 —
sparse 14.48 ± 0.42 34.5 ± 3.5 −100 —
elitist 88.90 ± 0.39 169.5 ± 10.5 −104 —
cem 76.07 ± 2.11 128 ± 0 −105.5 −73.1
fly 73.54 ± 16.25 158 ± 17 −103 −61.9

P1 confirmed at the margin (fly − cem = +30 finds vs one seed-sd 24, ddof 1): 1.25 sd, two seeds. P2 falsified (elitist > cem on both spaces). P3 falsified (fly's finds 11 meV shallower than CEM's). P4 ✓ (five arms in 75 min). The headline: the whole-brain head with a live gain field ties elitist on finds (158 ± 17 vs 169.5 ± 10.5) and loses to both elitist and CEM on AUC_Q — the rate of distinct finds per unit spend, which is the metric the generator line was built on (trial.py) — with the widest seed spread of any arm and the shallowest finds. On a reward that ranks, the fly's advantage over model-free search (E136, E185: flat reward and island surrogate) does not appear. What survives: the fly beats uniform and sparse by an order of magnitude and CEM on count. What does not: "more distinct qualifiers per unit of verifier cost than the baselines." Third seeds for fly and elitist (E195c) are queued to settle the count tie; the AUC gap (73.5 vs 88.9, four seed-sd of elitist) will not close with a third seed.

Related entries

Built with PRISMWebsite and visualizations made using Claude