Is the brain's own learning rule the one that makes the walkers do worse?
Partly. It lowered the mushroom body's score, but on the whole brain it raised it by 0.057; there the readout's rule did the harm.
In the log: Which learning rule lowers the ladder-p of the walk: the core's or the readout's?
mixedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 0 · energy model0 predictions · 1 result paragraphEXPERIMENTS.md lines 10711–10739, lines 10840–10860
What E165 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E165.svg).
Results
EXPERIMENTS.md · line 10840
E165 result — mixed, and it separates the two rules. Compartment readout (core KC→MBON
learning only), rung-0 p at every walker position, last-ten-round means:
core rule only mushroom off → on whole brain off → on
swarm mean p 0.153 → 0.107 (−0.046) 0.133 → 0.190 (+0.057)
top-four p 0.158 → 0.119 (−0.040) 0.076 → 0.053 (−0.023)
E162, both rules: 0.153 → 0.101 0.133 → 0.119
Prediction 1 confirmed: on the mushroom body the core rule alone lowers the swarm's p —
it is the core rule that hurts there. Prediction 2 falsified: on the whole brain the
core rule alone raises the swarm's p by +0.057, and it was the readout's delta rule that
pulled E162 down to 0.119. So the two substrates disagree about which rule helps, and the
readout rule — which E159 credited with the whole-brain gain — costs the whole-brain walk
0.07 in swarm p while (E161) it also carried that gain. Prediction 3's clean "both positive"
branch did not occur; the falsification condition (mushroom contribution ≥ +0.02) is not
met. Top-four p falls with learning on both substrates under both readouts — the
proposals chosen for verification get worse in ladder units over forty rounds — while E163's
two-hundred-round finds are the best any fly arm has had. That tension is recorded as open:
either the top-four metric over ten rounds is not what finds are made of, or the first
forty rounds are the bad ones. A two-hundred-round ladder-p pass is the experiment that
resolves it, and it is not run tonight.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 10711–10739
E165 — Which learning rule lowers the ladder-p of the walk: the core's or the readout's?
E162 measured lessons lowering the swarm's rung-0 p over forty rounds on both substrates
with the learned readout on. Two rules were active: MushroomBody.learn_head on KC→MBON,
and PopulationReadout's delta rule on the readout weights. Same four passes as E162, with
--readout mbon — the compartment readout, so only the core learns. Same seeds, rounds,
walkers, step, lr.
Predicted: (1) the mushroom body's learning contribution in p is still negative
(swarm mean over the last ten rounds lower with learning than without by at least 0.02) —
the core rule is the one that hurts, because it is the one E155 already showed relocating
walkers more with learning on; (2) the whole brain's is negative too, by less; (3) if
instead both contributions turn positive, the readout's delta rule was the culprit and the
core is cleared — that is the branch I consider less likely, since the readout starts from
the compartment weights and moves slowly.
Falsified if the mushroom contribution is positive by ≥ +0.02 — then the readout rule is
what damaged E162's walk, PopulationReadout gets its learning rate cut or its rule changed
to a proper least-squares update, and E159's learned-MBON gain has to be re-explained.
Wake-up collection — nothing new decided, everything accounted for. The E161 → E162 →
pytest chain finished: 252 passed, 2 skipped (1,393 s under contention; 249 plus the
three broadcast-encoder tests). E163's smoke pass ran end to end — encoder: broadcast onto 17,479 afferent nodes, two rounds, no finds, as a two-round smoke should — and the full
whole-brain broadcast run (200 x 4 x 2) started from it; prediction on record under E163.
E165 waits behind it. The Δ-check stands at 11 of 14 anchors; the six binaries and Hf,
Mo, Nb, Ta, Ti are in, and the residual on Ta/V/W needs the V and W elemental cells, which
are the last to run — so no partial residual is quoted. The rung-4 pw.x is still
converging. No new job was launched this hour: E163 is the one heavy job beside pw.x.
EXPERIMENTS.md · lines 10840–10860
E165 result — mixed, and it separates the two rules. Compartment readout (core KC→MBON
learning only), rung-0 p at every walker position, last-ten-round means:
core rule only mushroom off → on whole brain off → on
swarm mean p 0.153 → 0.107 (−0.046) 0.133 → 0.190 (+0.057)
top-four p 0.158 → 0.119 (−0.040) 0.076 → 0.053 (−0.023)
E162, both rules: 0.153 → 0.101 0.133 → 0.119
Prediction 1 confirmed: on the mushroom body the core rule alone lowers the swarm's p —
it is the core rule that hurts there. Prediction 2 falsified: on the whole brain the
core rule alone raises the swarm's p by +0.057, and it was the readout's delta rule that
pulled E162 down to 0.119. So the two substrates disagree about which rule helps, and the
readout rule — which E159 credited with the whole-brain gain — costs the whole-brain walk
0.07 in swarm p while (E161) it also carried that gain. Prediction 3's clean "both positive"
branch did not occur; the falsification condition (mushroom contribution ≥ +0.02) is not
met. Top-four p falls with learning on both substrates under both readouts — the
proposals chosen for verification get worse in ladder units over forty rounds — while E163's
two-hundred-round finds are the best any fly arm has had. That tension is recorded as open:
either the top-four metric over ten rounds is not what finds are made of, or the first
forty rounds are the bad ones. A two-hundred-round ladder-p pass is the experiment that
resolves it, and it is not run tonight.
Related entries
E162 — The walk measured in the ladder's own units, not the circuit's
E155 — The walk's own statistics on the two landscapes
E159 — Task 0, step 2: read the brain at LAL + DNa, not only at the 97 MBONs
E161 — Where the learned readout puts its weight (E159's unmeasured prediction 2)
E163 — Task 0, step 2b: the afferent broadcast as the brain's input