Does the fly beat a standard statistical model at choosing alloys for two goals?
No. The standard model scored 16.93, random 16.45 and the fly 15.73: a pre-registered loss, below random.
In the log: Pre-registration: the fly against a Gaussian process, several objectives
falsifiedDate not stated; the file was added to git on 2026-09-12 19:57generator · fly brain0 predictions · 0 result paragraphsPREREGISTER_multiobjective.md lines 1–52
What PREREGISTER:multiobjective did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/PREREGISTER-multiobjective.svg).
Pre-registration
Results
No result paragraph for this entry was found in the log.
The full record
PREREGISTER_multiobjective.md · lines 1–52
Pre-registration: the fly against a Gaussian process, several objectives
Written before the experiment is run, so the result cannot be chosen after the fact.
Committed to the repository at the same time as the code that runs it.
Hypothesis
With one objective a Gaussian process wins, measured: -481.8 +/- 3.0 against -470.5 +/- 4.3
meV/atom, t = 6.8 (E33). The stated reason is that the objective is smooth in eight
dimensions - the best case for a kernel - and the process carries a calibrated uncertainty
while the circuit's error is bias and cannot supply one (E29).
Hypothesis: with several objectives the gap narrows or reverses, because a mushroom
body is many compartments reading one shared odour code, each taught its own valence,
while a Gaussian process needs an independent model per objective and shares nothing
between them. The advantage should grow with the number of objectives.
This can fail, and the failure is informative: it would mean the architecture's
multi-compartment design buys nothing here either.
Design, fixed in advance
Objectives: free energy (minimise) and solid-solution fraction (maximise), both from
the validated environment at 128 sites. Their tension is measured, not assumed
(correlation +0.67, 16 of 200 on the front, E25).
Budget: 60 evaluations, 20 of them a Latin hypercube start. Same for every method.
Seeds: 10, indices 700-709, fixed now.
Metric: hypervolume of the Pareto front of true objective values among the points
evaluated, against a reference point fixed before the run from the 4,000-draw sweep -
the worst value of each objective in that sweep. Hypervolume is chosen because it is the
standard multi-objective measure and needs no weighting decided after seeing results.
Methods: random; a Gaussian process per objective with ParEGO scalarisation (the
standard multi-objective treatment, given the same 4,000-candidate pool it had in E33);
the fly with its compartments partitioned by measured valence, one group per objective,
sharing a single Kenyon code.
What counts as a win
Win: the fly's mean hypervolume exceeds the process's, with the difference at least
two standard errors, over the ten pre-registered seeds.
Draw: the difference is under two standard errors.
Loss: the process exceeds the fly by at least two standard errors.
What is not allowed
No seed selection, no re-running with different seeds and reporting the better set.
No change to the metric, reference point, budget or seed list after seeing any result.
No tuning of the fly on this benchmark. Sparsity stays at 0.12, carried over from E32,
and its own confirmation on an unseen pool remains outstanding and is reported as such.
The process keeps the generous treatment it had in E33: learned length scales, fitted
noise, and a candidate pool far larger than the fly ever proposes.
Every outcome gets recorded, including a loss, including a draw.
Files it names
PREREGISTER_multiobjective.md
Paths in the Forager repository, as the log wrote them.
Related entries
E33 — A Gaussian process beats the fly, and the reason is architectural
E29 — The fly's error is bias, so no variance-based uncertainty can see it
E25 — The thermodynamics wired in as a generator objective
E32 — The fly decides how long to forage, from the returns