No result paragraph for this entry was found in the log.
PREREGISTER_arena.md · lines 1–98Pre-registration: the fly steering a two-dimensional arena over composition space
Written before the arena has produced a single number. The thresholds below are decision
rules, not scientific constants, and they are deliberately demanding.
What is being tested
Does a controller read from the measured connectome steer a body through composition
space using fewer charged objective evaluations than the strongest small stateful
controller given exactly the same information?
Not "is the fly better than nothing". A memoryless rule that turns toward the better
antenna is trivial to beat with any filter, and beating it would establish nothing about
wiring.
What every controller receives, and nothing more
Two numbers per step: the contrast between the antennae, bounded and sign-symmetric so
it cannot leak where the objective's zero sits, and the level, how good it smells here
on a non-negative scale calibrated from the run's own history. The fly receives these as a
constant-sum drive split across the receptor cells of one glomerulus on each antenna; the
controls receive the same two numbers directly. Any controller may keep state.
Three objective evaluations are charged per walker per step - left antenna, right antenna,
body - and all three are banked: an alloy found by an antenna counts as found.
The arena is identical across controllers and is not part of the claim
Chart geometry, speed, turn limit, patience, relocation, restart policy and the minimum
element fraction belong to the harness. Chart planes are drawn from a stream keyed by
walker and relocation count, so the k-th plane of walker i is the same whichever
controller is steering. Re-anchoring a chart on the best alloy found is already an
objective-aware outer optimiser, and any advantage it confers belongs to the harness and
not to the animal; it is held fixed precisely so that it cancels.
Controls that must run
Functional calibration, before any benchmark
Passing these is a precondition, not a result. All are objective-free.
- No drift. With both antennae equal, the turn is zero after the measured bias is
subtracted.
- Correct sidedness. One-sided stimulation turns the body toward the stimulated side.
Already measured: DNa02 gives +2.50e-08 for left drive and -7.92e-08 for right, and
reverses; DNa01 does not reverse and is not used.
- Reversal under mirroring. Swapping antennae and steering cells flips the sign.
- A usable range. The turn varies across the contrast range rather than being ternary.
- Real movement. Composition displacement per step and antenna separation stay
non-degenerate; near a simplex face the projection can flatten the chart so that two
antennae sample the same alloy and the body stops moving in composition space. These
are counted and relocated out of, and reported.
Noise
The native evaluator at its real cost is the primary regime. No synthetic noise is added
on top of an already noisy evaluator, and no noise level is selected after seeing results.
The scatter of the returned evaluation is measured first, by repeated whole evaluations at
independently chosen compositions, and reported. The +/- 173 K figure for the
order-disorder transition is a separate uncertainty statement about a different quantity;
it is not a per-evaluation standard deviation and is not converted into one.
Kill criterion
The connectome-as-generator line is discontinued unless, in one confirmatory experiment on
held-out seeds after the design is frozen:
- it reaches an independently specified, re-verified quality target in at least 20%
fewer charged evaluations than the strongest pre-specified small stateful controller,
with the lower simultaneous 95% confidence bound above 20%; and
- it gives up no more than five percentage points of target-attainment probability,
also with a confidence bound; and
- it beats the matched shuffled-wiring ensemble.
Runs that fail to reach the target are assigned the full budget when timing is computed,
and attainment is reported separately; speed is never computed among winners only. The
returned alloy is re-evaluated at higher precision from a separate, equal budget, because
the largest noisy score seen is not a performance measure.
Results obtained only by selecting, after the fact, the noise level, the seeds, the
glomerulus, the readout cells or the relocation settings do not satisfy this criterion.
A failure means stop spending on this line and report the negative result. Where the
confidence interval is tight the finding is "no practically useful advantage in this
regime"; where it is wide the finding is "insufficient evidence", and the stop decision is
the same. Neither is a claim that every biological interface must fail.
Paths in the Forager repository, as the log wrote them.