Experiments · E37

Can the fly generate good new alloys that no pre-made list contained?

No. It generated alloys outside every list but gained 0.0123 of the score against the optimiser's 0.0750, a pre-registered loss.

In the log: The generative test: pre-registered, and a loss on both arms

falsifiedDate 2026-09-12, as written in the loggenerator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 1856–1902
exp E37 diagram
What E37 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E37.svg).

Pre-registration

The pre-registration, as written

E37 — The generative test: pre-registered, and a loss on both arms

Date 2026-09-12 · Question A Gaussian process cannot generate - it ranks a list it is handed. Can a learned generator produce crystals meeting several objectives, including compositions no list contained? · Provenance PREREGISTER_generative.md, …/generative.py

Why this differs from E35. That comparison had the circuit selecting from the same 4,000-point pool as the process, which removed the only structural difference between them. Here the circuit generates by run-and-tumble and the process does what it can do, which is rank a pool. The question standing behind it is whether a learned generator is viable at all - which is also the premise of the generative-flow-network work.

Result: LOSS, both arms. Hypervolume gained over the shared start, objectives normalised, matched greedy acquisition, seeds 800-809:

arm gained standard error left the pool
Gaussian process, selecting 0.0750 0.0117 90%*
random generation 0.0495 0.0085 30%*
the fly, generating 0.0123 0.0040 100%
the fly, generating + the expansion's uncertainty 0.0079 0.0032 100%

-5.09 and -5.56 standard errors. The circuit generates compositions outside any pool on every seed and they are not good ones. Adding the expansion's measured uncertainty as an exploration term made it slightly worse, not better.

A flaw in one column, found after the run and reported rather than dropped. The process was given its own Latin hypercube pool (a different draw from the reference), so "left the pool" measures only that its draw contained points the reference draw did not dominate - not that it generated anything, which it cannot. That column distinguishes nothing as built; to mean what it was meant to, the process's pool must be the reference pool. The hypervolume comparison is unaffected.

Also settled, separately: the promotion signal changes nothing. Teaching the circuit a verdict - was this candidate worth its promotion, +1 or -1 - rather than an objective's value is the better fit for machinery built on appetitive against aversive dopamine, and was argued for on those grounds before testing. Measured: 0.7737 +/- 0.0184 against 0.7783 +/- 0.0142 for the value signal. No difference.

Standing count. Five fair configurations now - single-objective selection (E33), representation substitution (E33), multi-objective selection (E35), generative multi-objective (here), and promotion-based learning (here). The circuit has won none of them. Each loss has had a mechanism found and fixed, and the next has followed. That pattern is itself the finding: repeated repair with the outcome fixed in advance is how noise gets fitted, and a favourable result arriving after enough of it would not deserve belief.


Results

EXPERIMENTS.md · line 1869

Result: LOSS, both arms. Hypervolume gained over the shared start, objectives normalised, matched greedy acquisition, seeds 800-809:

The full record

EXPERIMENTS.md · lines 1856–1902

E37 — The generative test: pre-registered, and a loss on both arms

Date 2026-09-12 · Question A Gaussian process cannot generate - it ranks a list it is handed. Can a learned generator produce crystals meeting several objectives, including compositions no list contained? · Provenance PREREGISTER_generative.md, …/generative.py

Why this differs from E35. That comparison had the circuit selecting from the same 4,000-point pool as the process, which removed the only structural difference between them. Here the circuit generates by run-and-tumble and the process does what it can do, which is rank a pool. The question standing behind it is whether a learned generator is viable at all - which is also the premise of the generative-flow-network work.

Result: LOSS, both arms. Hypervolume gained over the shared start, objectives normalised, matched greedy acquisition, seeds 800-809:

arm gained standard error left the pool
Gaussian process, selecting 0.0750 0.0117 90%*
random generation 0.0495 0.0085 30%*
the fly, generating 0.0123 0.0040 100%
the fly, generating + the expansion's uncertainty 0.0079 0.0032 100%

-5.09 and -5.56 standard errors. The circuit generates compositions outside any pool on every seed and they are not good ones. Adding the expansion's measured uncertainty as an exploration term made it slightly worse, not better.

A flaw in one column, found after the run and reported rather than dropped. The process was given its own Latin hypercube pool (a different draw from the reference), so "left the pool" measures only that its draw contained points the reference draw did not dominate - not that it generated anything, which it cannot. That column distinguishes nothing as built; to mean what it was meant to, the process's pool must be the reference pool. The hypervolume comparison is unaffected.

Also settled, separately: the promotion signal changes nothing. Teaching the circuit a verdict - was this candidate worth its promotion, +1 or -1 - rather than an objective's value is the better fit for machinery built on appetitive against aversive dopamine, and was argued for on those grounds before testing. Measured: 0.7737 +/- 0.0184 against 0.7783 +/- 0.0142 for the value signal. No difference.

Standing count. Five fair configurations now - single-objective selection (E33), representation substitution (E33), multi-objective selection (E35), generative multi-objective (here), and promotion-based learning (here). The circuit has won none of them. Each loss has had a mechanism found and fixed, and the next has followed. That pattern is itself the finding: repeated repair with the outcome fixed in advance is how noise gets fitted, and a favourable result arriving after enough of it would not deserve belief.

Related entries

Built with PRISMWebsite and visualizations made using Claude