Experiments · E44

Are the fly's proposals worse than evenly spread ones when only the proposer changes?

No. The same optimiser reached −479.9 from the fly's proposals and −479.0 meV/atom from evenly spread ones: a draw, 1.2 standard errors.

In the log: With the confound removed, the circuit's proposals match space-filling

recordedDate 2026-09-13, as written in the logunclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 2234–2275
exp E44 diagram
What E44 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E44.svg).

Results

No result paragraph for this entry was found in the log.

The full record

EXPERIMENTS.md · lines 2234–2275

E44 — With the confound removed, the circuit's proposals match space-filling

Date 2026-09-13 · Question (operator, twice) A Gaussian process cannot generate. Comparing it to the circuit compares two things at once. What happens when only the generator varies? · Provenance …/generator_arm.py, METHODOLOGY 4i

The confound, stated plainly. E33, E35 and E37 all reported the circuit against a Gaussian process. A process cannot propose a composition - it ranks a list - so in every one of those the list came from a 4,000-point Latin hypercube. What was measured was (hypercube generating + process selecting) against (circuit generating + circuit selecting), two variables at once, and written up as though it were one.

The arm that isolates the generator. Same process, same expected-improvement acquisition, same budget of 60, same 1,500-candidate pool each round. The only difference is who produced the candidates. Eight seeds.

who proposes best F found standard error
Latin hypercube, space-filling -479.0 +/- 1.3 0.5
the circuit, walking and proposing -479.9 +/- 1.6 0.6

-0.9 +/- 0.7 meV/atom, 1.2 standard errors: a draw. The circuit's proposals are as good as a space-filling design, and no better. The direction favours the circuit and the effect is not established - resolving a difference this small at two standard errors would need roughly twenty seeds.

This is the first comparison the circuit has not lost, and the reason is that it is the first one asking a single question. The earlier margins - 11.3 meV/atom in E33, five standard errors in E35 and E37 - were substantially the confound rather than the circuit.

An artefact, flagged so it is not mistaken for a result. The circuit's arm ran in 561 s against the hypercube's 7,146 s. That is HEASpace.sample doing rejection sampling at a 200-fold oversample, not anything about the circuit. Timing from this experiment means nothing.

What it does and does not establish. It does not show a learned generator beats space-filling sampling. It does show the two are within noise of each other on this problem, and it removes the basis for "the fly lost to a Gaussian process", which was never a claim this project had earned. More seeds would settle the direction; the design space here is eight-dimensional and bounded, which is exactly where a space-filling design is at its strongest, so a draw in that setting is a stronger result for the circuit than it first appears.

Related entries

Built with PRISMWebsite and visualizations made using Claude