Are the fly's proposals worse than evenly spread ones when only the proposer changes?
No. The same optimiser reached −479.9 from the fly's proposals and −479.0 meV/atom from evenly spread ones: a draw, 1.2 standard errors.
In the log: With the confound removed, the circuit's proposals match space-filling
recordedDate 2026-09-13, as written in the logunclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 2234–2275
What E44 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E44.svg).
Results
No result paragraph for this entry was found in the log.
The full record
EXPERIMENTS.md · lines 2234–2275
E44 — With the confound removed, the circuit's proposals match space-filling
Date 2026-09-13 · Question (operator, twice) A Gaussian process cannot generate.
Comparing it to the circuit compares two things at once. What happens when only the
generator varies? · Provenance…/generator_arm.py, METHODOLOGY 4i
The confound, stated plainly.E33, E35 and E37 all reported the circuit against a
Gaussian process. A process cannot propose a composition - it ranks a list - so in every
one of those the list came from a 4,000-point Latin hypercube. What was measured was
(hypercube generating + process selecting) against (circuit generating + circuit
selecting), two variables at once, and written up as though it were one.
The arm that isolates the generator. Same process, same expected-improvement
acquisition, same budget of 60, same 1,500-candidate pool each round. The only difference
is who produced the candidates. Eight seeds.
who proposes
best F found
standard error
Latin hypercube, space-filling
-479.0 +/- 1.3
0.5
the circuit, walking and proposing
-479.9 +/- 1.6
0.6
-0.9 +/- 0.7 meV/atom, 1.2 standard errors: a draw. The circuit's proposals are as good
as a space-filling design, and no better. The direction favours the circuit and the effect
is not established - resolving a difference this small at two standard errors would need
roughly twenty seeds.
This is the first comparison the circuit has not lost, and the reason is that it is the
first one asking a single question. The earlier margins - 11.3 meV/atom in E33, five
standard errors in E35 and E37 - were substantially the confound rather than the circuit.
An artefact, flagged so it is not mistaken for a result. The circuit's arm ran in 561 s
against the hypercube's 7,146 s. That is HEASpace.sample doing rejection sampling at a
200-fold oversample, not anything about the circuit. Timing from this experiment means
nothing.
What it does and does not establish. It does not show a learned generator beats
space-filling sampling. It does show the two are within noise of each other on this
problem, and it removes the basis for "the fly lost to a Gaussian process", which was never
a claim this project had earned. More seeds would settle the direction; the design space
here is eight-dimensional and bounded, which is exactly where a space-filling design is at
its strongest, so a draw in that setting is a stronger result for the circuit than it
first appears.
Related entries
E33 — A Gaussian process beats the fly, and the reason is architectural
E35 — The multi-objective test: a pre-registered loss, on a benchmark with no power
E37 — The generative test: pre-registered, and a loss on both arms