Can the fly generate good new alloys that no pre-made list contained?
No. It generated alloys outside every list but gained 0.0123 of the score against the optimiser's 0.0750, a pre-registered loss.
In the log: The generative test: pre-registered, and a loss on both arms
falsifiedDate 2026-09-12, as written in the loggenerator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 1856–1902
What E37 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E37.svg).
Pre-registration
The pre-registration, as written
E37 — The generative test: pre-registered, and a loss on both arms
Date 2026-09-12 · Question A Gaussian process cannot generate - it ranks a list it
is handed. Can a learned generator produce crystals meeting several objectives, including
compositions no list contained? · ProvenancePREREGISTER_generative.md,
…/generative.py
Why this differs from E35. That comparison had the circuit selecting from the same
4,000-point pool as the process, which removed the only structural difference between them.
Here the circuit generates by run-and-tumble and the process does what it can do, which is
rank a pool. The question standing behind it is whether a learned generator is viable at
all - which is also the premise of the generative-flow-network work.
Result: LOSS, both arms. Hypervolume gained over the shared start, objectives
normalised, matched greedy acquisition, seeds 800-809:
arm
gained
standard error
left the pool
Gaussian process, selecting
0.0750
0.0117
90%*
random generation
0.0495
0.0085
30%*
the fly, generating
0.0123
0.0040
100%
the fly, generating + the expansion's uncertainty
0.0079
0.0032
100%
-5.09 and -5.56 standard errors. The circuit generates compositions outside any pool on
every seed and they are not good ones. Adding the expansion's measured uncertainty as an
exploration term made it slightly worse, not better.
A flaw in one column, found after the run and reported rather than dropped. The process
was given its own Latin hypercube pool (a different draw from the reference), so "left the
pool" measures only that its draw contained points the reference draw did not dominate - not
that it generated anything, which it cannot. That column distinguishes nothing as built;
to mean what it was meant to, the process's pool must be the reference pool. The
hypervolume comparison is unaffected.
Also settled, separately: the promotion signal changes nothing. Teaching the circuit a
verdict - was this candidate worth its promotion, +1 or -1 - rather than an objective's
value is the better fit for machinery built on appetitive against aversive dopamine, and
was argued for on those grounds before testing. Measured: 0.7737 +/- 0.0184 against 0.7783
+/- 0.0142 for the value signal. No difference.
Standing count. Five fair configurations now - single-objective selection (E33),
representation substitution (E33), multi-objective selection (E35), generative
multi-objective (here), and promotion-based learning (here). The circuit has won none of
them. Each loss has had a mechanism found and fixed, and the next has followed. That
pattern is itself the finding: repeated repair with the outcome fixed in advance is how
noise gets fitted, and a favourable result arriving after enough of it would not deserve
belief.
Results
EXPERIMENTS.md · line 1869
Result: LOSS, both arms. Hypervolume gained over the shared start, objectives
normalised, matched greedy acquisition, seeds 800-809:
The full record
EXPERIMENTS.md · lines 1856–1902
E37 — The generative test: pre-registered, and a loss on both arms
Date 2026-09-12 · Question A Gaussian process cannot generate - it ranks a list it
is handed. Can a learned generator produce crystals meeting several objectives, including
compositions no list contained? · ProvenancePREREGISTER_generative.md,
…/generative.py
Why this differs from E35. That comparison had the circuit selecting from the same
4,000-point pool as the process, which removed the only structural difference between them.
Here the circuit generates by run-and-tumble and the process does what it can do, which is
rank a pool. The question standing behind it is whether a learned generator is viable at
all - which is also the premise of the generative-flow-network work.
Result: LOSS, both arms. Hypervolume gained over the shared start, objectives
normalised, matched greedy acquisition, seeds 800-809:
arm
gained
standard error
left the pool
Gaussian process, selecting
0.0750
0.0117
90%*
random generation
0.0495
0.0085
30%*
the fly, generating
0.0123
0.0040
100%
the fly, generating + the expansion's uncertainty
0.0079
0.0032
100%
-5.09 and -5.56 standard errors. The circuit generates compositions outside any pool on
every seed and they are not good ones. Adding the expansion's measured uncertainty as an
exploration term made it slightly worse, not better.
A flaw in one column, found after the run and reported rather than dropped. The process
was given its own Latin hypercube pool (a different draw from the reference), so "left the
pool" measures only that its draw contained points the reference draw did not dominate - not
that it generated anything, which it cannot. That column distinguishes nothing as built;
to mean what it was meant to, the process's pool must be the reference pool. The
hypervolume comparison is unaffected.
Also settled, separately: the promotion signal changes nothing. Teaching the circuit a
verdict - was this candidate worth its promotion, +1 or -1 - rather than an objective's
value is the better fit for machinery built on appetitive against aversive dopamine, and
was argued for on those grounds before testing. Measured: 0.7737 +/- 0.0184 against 0.7783
+/- 0.0142 for the value signal. No difference.
Standing count. Five fair configurations now - single-objective selection (E33),
representation substitution (E33), multi-objective selection (E35), generative
multi-objective (here), and promotion-based learning (here). The circuit has won none of
them. Each loss has had a mechanism found and fixed, and the next has followed. That
pattern is itself the finding: repeated repair with the outcome fixed in advance is how
noise gets fitted, and a favourable result arriving after enough of it would not deserve
belief.
Related entries
E35 — The multi-objective test: a pre-registered loss, on a benchmark with no power
E33 — A Gaussian process beats the fly, and the reason is architectural