Experiments · E205

Do all three generator fixes together let the fly beat keep-the-best-and-mutate?

No. The fly found 149 alloys against elitist's 238, and more slowly; two of the three fixes had already failed alone.

In the log: E202–E205 — the three fixes, one at a time, then together (2026-09-19 08:47; queued after E201 and E173c)

falsifiedDate 2026-09-19 08:47, as written in the loggenerator · fly brain0 predictions · 0 result paragraphsEXPERIMENTS.md lines 12676–12695, lines 12955–12958
exp E205 diagram
What E205 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E205.svg).

Pre-registration

The pre-registration, as written

E205 — all three fixes together, 3 seeds, own elitist. fly 81.9 ± 10.5 (s.e.) / 149 finds vs elitist 144.8 ± 6.3 / 238. Falsified on rate and finds, confirmed on depth. As stated before it ran: void as specified — it combined two falsified changes, and the combination is worse than either alone.

Results

No result paragraph for this entry was found in the log.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 12676–12695

E202–E205 — the three fixes, one at a time, then together (2026-09-19 08:47; queued after E201 and E173c)

Correction to the E195/E195b tables above: the log's ± is the standard error over seeds (sd/√n), not the population sd; score_arms.py reports the ddof-1 sd. Baseline = E195b fly: AUC_Q 73.54 (s.e. 16.25, sd 23), 158 finds (sd 24), last-window finds 123 / 113, mean find depth −61.9; elitist 88.90 / 169.5 / last-window 147, 139 / depth −74.4 (seed mean). Code: branch fly-competitive (local search in Forager, imagination + history in FlyGenerator, prediction_error in mushroom.py, flags + diagnostics in stage_b; flags off reproduce E195b bit-for-bit by construction and by test). Predictions, each before its run:

  • E202 local search (FORAGER_LOCAL_SEARCH=1): last-window finds ≥ 139 (elitist's floor) and distinct finds ≥ 170; AUC_Q within one seed-sd of elitist. If the last window does not move, the stall is not the walkers' memory and the anchor mechanism is wrong as built.
  • E203 imagination (FORAGER_IMAGINE=4): mean find depth deeper by ≥ 5 meV than −61.9; finds unchanged within sd. Void if E201's score–reward correlation is < 0.2.
  • E204 RPE (FORAGER_RPE=bennett): score–reward correlation up by ≥ 0.10 over E201's live arm; AUC_Q not worse.
  • E205 all three vs elitist, 3 seeds: fly AUC_Q ≥ elitist − one seed-sd, finds ≥ elitist, depth within 5 meV of elitist's. If E205 fails after E202 passed, the pieces interfere and the combination is retuned one flag at a time, not all together.
EXPERIMENTS.md · lines 12955–12958

E205 — all three fixes together, 3 seeds, own elitist. fly 81.9 ± 10.5 (s.e.) / 149 finds vs elitist 144.8 ± 6.3 / 238. Falsified on rate and finds, confirmed on depth. As stated before it ran: void as specified — it combined two falsified changes, and the combination is worse than either alone.

Related entries

Built with PRISMWebsite and visualizations made using Claude