Experiments · E195

On the new reward, does the fly brain beat simple searches that use no brain?

Withdrawn. Keep-the-best-and-mutate matched or beat the fly here, but that reading reversed once a distorting gate was removed from the reward.

In the log: the fly against the model-free arms on the real reward (2026-09-18 22:58, N1)

supersededDate 2026-09-18 22:58, as written in the logrung 4 · DFT2 predictions · 1 result paragraphEXPERIMENTS.md lines 12268–12282, lines 12284–12294, lines 12296–12309, lines 12311–12314, lines 12495–12527
exp E195 diagram
What E195 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E195.svg).

Pre-registration

  1. (10)
    the last inequality of prediction 2 ("elitist < CEM") is falsified on the 12-element space; CEM's Dirichlet updates cannot recover from a batch paid all zeros, elitist's mutation of a kept draw can. fly pending.
    no verdict written against it
  2. (27)
    > cem
    no verdict written against it
The pre-registration, as written

E195 — the fly against the model-free arms on the real reward (2026-09-18 22:58, N1)

Every fly-vs-CEM number in this record (E136 13.6 / 34.5; E167 25.9 / 71; E171b 40.0 / 82) was earned on the flat icet reward, where every find scored ≈ −11 and the verdict could not separate good from better. On FORAGER_RUNG0=ece the reward ranks (E193). So the comparison is run again from scratch on the real reward: uniform (the null), sparse (Dirichlet α 0.3 — the fly's geometry without the fly), elitist, CEM (the strongest model-free arm), and the fly (E171b's head, MPS); 200 × 4 × 2 each, one process, tag e195. Predictions. 1. The fly beats CEM on finds by more than one seed-sd (E193's 61 / 77 against CEM's number here); if it does not, the head's contribution on a reward that ranks is null and every prior fly-vs-CEM claim is confined to the flat reward. 2. The model-free order is uniform < sparse < elitist < CEM, as on the flat reward (E136). 3. The fly's finds have a lower mean v5 energy than CEM's (−10 meV or more): a live gain field should find deeper minima, not only more of them. 4. Wall time: the CPU arms ~10 min each, the fly ~35 min; all five inside 1.5 h.

E195 confound, seen at its first two arms (23:05) and stated before the fly's number lands. uniform: 0 finds; sparse: 0 finds, best 0. On the 12-element space nearly every Dirichlet draw carries Co, Cu or Ni; the eCE has never seen them and the v5 screen refuses such compositions outright (p = 0, ece_foreign_fraction). The fly's walkers can drive a fraction to exactly zero (min_fraction 0) and so escape the refusal; a CEM or elitist arm paid zero for everything cannot learn where to go. So E195 on the 12-element space would give the fly a win it did not earn. E195b, queued behind it: the same five arms on the eight-element refractory space (FORAGER_CE=ce_8element: Hf Mo Nb Ta Ti V W Zr — all inside the eCE's nine, nothing refused), same predictions. E195's fly arm still stands as the 12-element reference; its model-free arms are void by construction, and are recorded as such, not as a result.

E195, third arm (23:07) — the confound note above was too strong. elitist: AUC_Q 27.08 ± 5.12, distinct 57 ± 14, best −103 meV/atom. So a model-free arm that learns does escape the foreign-element refusal (it keeps the rare all-refractory draws and mutates around them); only the two non-learning arms (uniform, sparse) are void on the 12-element space. Prediction 2 ("uniform < sparse < elitist < CEM") so far reads 0 = 0 < 27 — the first inequality is a tie at zero, not an ordering. Predictions 1 and 3 still wait on cem and fly. E195b (8-element) remains the clean comparison and is queued.

Convention drift caught before launch (23:10). The eight E192 cells and E194 were written with a 5×5×5 mesh; every validated 54-atom standard cell (E172b, E183, E187 — the ones within 10 meV of v5, and the rows v6 will train on) is 60/720 Ry at 4×4×4 (36 irreducible k-points; 5×5×5 is 63, ~1.75× the time and ~20 GB scratch per cell). All nine inputs corrected to 4×4×4 before pw.x reached them; nothing was restarted. Rule restated: the 54-atom standard is 60/720, k 4×4×4, MV 0.02 Ry, local-TF β 0.1.

E195, fourth arm (23:12). cem: AUC_Q 10.15 ± 10.15, distinct 21 ± 21, best −56 — one seed found nothing at all. elitist (27) > cem (10): the last inequality of prediction 2 ("elitist < CEM") is falsified on the 12-element space; CEM's Dirichlet updates cannot recover from a batch paid all zeros, elitist's mutation of a kept draw can. fly pending.

Results

EXPERIMENTS.md · line 12495

E195 result (23:51) — five arms on the v5 reward, 12-element space, 200 × 4 × 2 seeds.

arm AUC_Q distinct finds best v5 mean v5 of finds
uniform 0.00 0 +83 —
sparse 0.00 0 0 —
elitist 27.08 ± 5.12 57 ± 14 −103.7 —
cem 10.15 ± 10.15 21 ± 21 −90.5 −70.3
fly 20.68 ± 8.38 54.5 ± 16.5 −102.4 −62.2

(± in the table = population sd of two seeds, as the log prints; score_arms.py uses ddof = 1.) P1 confirmed as written (fly − cem = +33.5 finds > one seed-sd 29.7) but void in substance — CEM was starved by the refusal confound. P2 falsified (elitist 57

cem 21). P3 falsified: the fly's finds are shallower than CEM's by 8 meV. P4 confirmed (all five in 53 min). The headline is the arm the confound did not touch: elitist ≥ fly on every column (AUC 27.1 vs 20.7, finds 57 vs 54.5, best −104 vs −102). On a reward that ranks, the whole-brain head with a live gain field does no better than keep-the-best-and-mutate. E195b (eight elements, nothing refused) is the clean version and is queued; if elitist ≥ fly holds there, the fly's advantage was a property of the flat reward and the island surrogate, and the generator line's claim rests on nothing.

Search queue collapsed and rebuilt (23:56). Two crashes, both mine. (1) E173c died at its first confirm: confirm() taught the "5ht" site by name on a head built without one (StopIteration) — the run was the no-5-HT control, so the crash was the design. Guarded: a head teaches the site only if it carries it; test added. (2) E195b, then E198's two arms, died in build(): FORAGER_RUNG0=ece demanded that the expansion contain every eCE element, so the eight-element space (no Cr) was refused although the verifier projects by name. Relaxed to "shares at least one element" (ece_covers, tested). Each crashed chain still wrote its done-marker, so the queue fell through in seconds and E199 was killed before it could add a third. False markers removed, crashed logs parked in runs/crashed_2309/, and the four searches relaunched as one chain (runs/night_searches_chain.sh: E195b → E198 → E199 → E173c), one search at a time, load-gated. Predictions for all four stand as written; the v6 chain waits on the last marker.

The full record

This entry is written in 5 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 12268–12282

E195 — the fly against the model-free arms on the real reward (2026-09-18 22:58, N1)

Every fly-vs-CEM number in this record (E136 13.6 / 34.5; E167 25.9 / 71; E171b 40.0 / 82) was earned on the flat icet reward, where every find scored ≈ −11 and the verdict could not separate good from better. On FORAGER_RUNG0=ece the reward ranks (E193). So the comparison is run again from scratch on the real reward: uniform (the null), sparse (Dirichlet α 0.3 — the fly's geometry without the fly), elitist, CEM (the strongest model-free arm), and the fly (E171b's head, MPS); 200 × 4 × 2 each, one process, tag e195. Predictions. 1. The fly beats CEM on finds by more than one seed-sd (E193's 61 / 77 against CEM's number here); if it does not, the head's contribution on a reward that ranks is null and every prior fly-vs-CEM claim is confined to the flat reward. 2. The model-free order is uniform < sparse < elitist < CEM, as on the flat reward (E136). 3. The fly's finds have a lower mean v5 energy than CEM's (−10 meV or more): a live gain field should find deeper minima, not only more of them. 4. Wall time: the CPU arms ~10 min each, the fly ~35 min; all five inside 1.5 h.

EXPERIMENTS.md · lines 12284–12294

E195 confound, seen at its first two arms (23:05) and stated before the fly's number lands. uniform: 0 finds; sparse: 0 finds, best 0. On the 12-element space nearly every Dirichlet draw carries Co, Cu or Ni; the eCE has never seen them and the v5 screen refuses such compositions outright (p = 0, ece_foreign_fraction). The fly's walkers can drive a fraction to exactly zero (min_fraction 0) and so escape the refusal; a CEM or elitist arm paid zero for everything cannot learn where to go. So E195 on the 12-element space would give the fly a win it did not earn. E195b, queued behind it: the same five arms on the eight-element refractory space (FORAGER_CE=ce_8element: Hf Mo Nb Ta Ti V W Zr — all inside the eCE's nine, nothing refused), same predictions. E195's fly arm still stands as the 12-element reference; its model-free arms are void by construction, and are recorded as such, not as a result.

EXPERIMENTS.md · lines 12296–12309

E195, third arm (23:07) — the confound note above was too strong. elitist: AUC_Q 27.08 ± 5.12, distinct 57 ± 14, best −103 meV/atom. So a model-free arm that learns does escape the foreign-element refusal (it keeps the rare all-refractory draws and mutates around them); only the two non-learning arms (uniform, sparse) are void on the 12-element space. Prediction 2 ("uniform < sparse < elitist < CEM") so far reads 0 = 0 < 27 — the first inequality is a tie at zero, not an ordering. Predictions 1 and 3 still wait on cem and fly. E195b (8-element) remains the clean comparison and is queued.

Convention drift caught before launch (23:10). The eight E192 cells and E194 were written with a 5×5×5 mesh; every validated 54-atom standard cell (E172b, E183, E187 — the ones within 10 meV of v5, and the rows v6 will train on) is 60/720 Ry at 4×4×4 (36 irreducible k-points; 5×5×5 is 63, ~1.75× the time and ~20 GB scratch per cell). All nine inputs corrected to 4×4×4 before pw.x reached them; nothing was restarted. Rule restated: the 54-atom standard is 60/720, k 4×4×4, MV 0.02 Ry, local-TF β 0.1.

EXPERIMENTS.md · lines 12311–12314

E195, fourth arm (23:12). cem: AUC_Q 10.15 ± 10.15, distinct 21 ± 21, best −56 — one seed found nothing at all. elitist (27) > cem (10): the last inequality of prediction 2 ("elitist < CEM") is falsified on the 12-element space; CEM's Dirichlet updates cannot recover from a batch paid all zeros, elitist's mutation of a kept draw can. fly pending.

EXPERIMENTS.md · lines 12495–12527

E195 result (23:51) — five arms on the v5 reward, 12-element space, 200 × 4 × 2 seeds.

arm AUC_Q distinct finds best v5 mean v5 of finds
uniform 0.00 0 +83 —
sparse 0.00 0 0 —
elitist 27.08 ± 5.12 57 ± 14 −103.7 —
cem 10.15 ± 10.15 21 ± 21 −90.5 −70.3
fly 20.68 ± 8.38 54.5 ± 16.5 −102.4 −62.2

(± in the table = population sd of two seeds, as the log prints; score_arms.py uses ddof = 1.) P1 confirmed as written (fly − cem = +33.5 finds > one seed-sd 29.7) but void in substance — CEM was starved by the refusal confound. P2 falsified (elitist 57

cem 21). P3 falsified: the fly's finds are shallower than CEM's by 8 meV. P4 confirmed (all five in 53 min). The headline is the arm the confound did not touch: elitist ≥ fly on every column (AUC 27.1 vs 20.7, finds 57 vs 54.5, best −104 vs −102). On a reward that ranks, the whole-brain head with a live gain field does no better than keep-the-best-and-mutate. E195b (eight elements, nothing refused) is the clean version and is queued; if elitist ≥ fly holds there, the fly's advantage was a property of the flat reward and the island surrogate, and the generator line's claim rests on nothing.

Search queue collapsed and rebuilt (23:56). Two crashes, both mine. (1) E173c died at its first confirm: confirm() taught the "5ht" site by name on a head built without one (StopIteration) — the run was the no-5-HT control, so the crash was the design. Guarded: a head teaches the site only if it carries it; test added. (2) E195b, then E198's two arms, died in build(): FORAGER_RUNG0=ece demanded that the expansion contain every eCE element, so the eight-element space (no Cr) was refused although the verifier projects by name. Relaxed to "shares at least one element" (ece_covers, tested). Each crashed chain still wrote its done-marker, so the queue fell through in seconds and E199 was killed before it could add a third. False markers removed, crashed logs parked in runs/crashed_2309/, and the four searches relaunched as one chain (runs/night_searches_chain.sh: E195b → E198 → E199 → E173c), one search at a time, load-gated. Predictions for all four stand as written; the v6 chain waits on the last marker.

Related entries

Built with PRISMWebsite and visualizations made using Claude