Measured by the ladder's own verdict, do lessons move the walkers toward better alloys?
No. With lessons on, the swarm's average ladder score fell on both brains, from 0.153 to 0.101 on the mushroom body.
In the log: The walk measured in the ladder's own units, not the circuit's
falsifiedDate not stated in the log; it was written between the commit of 2026-09-16 19:06 and the first commit that contains it, 2026-09-19 08:35rung 0 · energy model0 predictions · 1 result paragraphEXPERIMENTS.md lines 10518–10542, lines 10684–10707
What E162 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E162.svg).
Results
EXPERIMENTS.md · line 10684
E162 result — falsified on all three counts, and the pre-registered branch fires.
Rung-0 p at every walker position, mean over the last ten of forty rounds:
learned readout mushroom off → on whole brain off → on
swarm mean p 0.153 → 0.101 (−0.052) 0.133 → 0.119 (−0.014)
top-four p 0.158 → 0.134 (−0.024) 0.076 → 0.072 (−0.004)
first-ten swarm p 0.077 0.077 0.107 0.106
Prediction 1 falsified (mushroom gain ≥ +0.05: it is −0.05). Prediction 2 falsified
(whole-brain sign positive: −0.014). Prediction 3 falsified (top-four p rises with
learning: it falls on both). The falsification condition is met, so: the walk explains
none of E159's 17x, and E160's sign was not the ruler. The greedy walk alone doubles the
swarm's p over forty rounds on the mushroom body (0.077 → 0.153) — the circuit's initial
landscape correlates with the ladder — and lessons make it climb less, on both substrates,
measured in the units the search is paid in. Finds in E159 therefore do not come from the
swarm rising; they come from the round's top four being verified out of a swarm that is not
improving. Two learning rules were on at once — the core's KC→MBON plasticity and the
readout's delta rule — and this cannot say which one hurts.
Second reading, stated so it is tested rather than absorbed: forty rounds is ten times
fewer lessons than E159's two hundred, and the first two lessons teach nothing; a rule that
overshoots early and settles late would look exactly like this. E165 runs forty rounds with
the compartment readout to remove one rule; a two-hundred-round ladder-p pass is the follow-up
if E165 clears the core.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 10518–10542
E162 — The walk measured in the ladder's own units, not the circuit's
E160's climb metric divides by a landscape sd fixed before learning; a learned readout
rescales the score as its weights grow, so "climb in sd" is not scale-free and E160's
−0.58 for the whole brain may be an artefact of the ruler. The scale-free measure is the
ladder's own verdict at each walker's position: Verifier.screen mapped to p exactly as
the search's reward is, evaluated on all 24 walkers every round. Same four passes as E160
(both assets, learned readout, learning off then on), same seeds; recorded per round: mean
p over the swarm, and mean p over the round's top four by circuit score — the proposals the
search would actually send for verification.
Predicted:
The mushroom body's learning contribution is positive in p: mean swarm p over the
last ten rounds is at least +0.05 higher with learning on than off.
The whole brain's sign flips back: its learning contribution in p is positive
(≥ +0.03), which would make E160's −0.58 the rescaling artefact I suspected.
Top-four p rises with learning on both assets, and the whole brain's top-four gain is
at least the mushroom body's — consistent with E159's 7 finds against 5.5: the readout
improves what gets verified more than how far the swarm climbs.
Falsified if the whole brain's top-four p with learning on is no higher than with it off.
Then the walk explains none of E159's 17x, the metric was not the problem, and what is left
is verification-time selection or two lucky seeds — in which case E161's second pair of
seeds, with the readout saved, is the arbiter and step 2b should not be built on E159 alone.
EXPERIMENTS.md · lines 10684–10707
E162 result — falsified on all three counts, and the pre-registered branch fires.
Rung-0 p at every walker position, mean over the last ten of forty rounds:
learned readout mushroom off → on whole brain off → on
swarm mean p 0.153 → 0.101 (−0.052) 0.133 → 0.119 (−0.014)
top-four p 0.158 → 0.134 (−0.024) 0.076 → 0.072 (−0.004)
first-ten swarm p 0.077 0.077 0.107 0.106
Prediction 1 falsified (mushroom gain ≥ +0.05: it is −0.05). Prediction 2 falsified
(whole-brain sign positive: −0.014). Prediction 3 falsified (top-four p rises with
learning: it falls on both). The falsification condition is met, so: the walk explains
none of E159's 17x, and E160's sign was not the ruler. The greedy walk alone doubles the
swarm's p over forty rounds on the mushroom body (0.077 → 0.153) — the circuit's initial
landscape correlates with the ladder — and lessons make it climb less, on both substrates,
measured in the units the search is paid in. Finds in E159 therefore do not come from the
swarm rising; they come from the round's top four being verified out of a swarm that is not
improving. Two learning rules were on at once — the core's KC→MBON plasticity and the
readout's delta rule — and this cannot say which one hurts.
Second reading, stated so it is tested rather than absorbed: forty rounds is ten times
fewer lessons than E159's two hundred, and the first two lessons teach nothing; a rule that
overshoots early and settles late would look exactly like this. E165 runs forty rounds with
the compartment readout to remove one rule; a two-hundred-round ladder-p pass is the follow-up
if E165 clears the core.
Related entries
E160 — E155's walk statistics, re-run under the learned LAL + DNa + MBON readout
E159 — Task 0, step 2: read the brain at LAL + DNa, not only at the 97 MBONs
E161 — Where the learned readout puts its weight (E159's unmeasured prediction 2)
E165 — Which learning rule lowers the ladder-p of the walk: the core's or the readout's?