Once the model's vanadium error stops paying, what does the site's generator run find?
Partly. Vanadium-rich picks vanished and the family stayed Mo–Nb–Ta–W, but the best score was −90 meV/atom, short of the predicted −110 to −120.
In the log: Re-recording the site run: prediction before the rerun
mixedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04generator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 4276–4344
What E79 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E79.svg).
Pre-registration
The pre-registration, as written
E79 — Re-recording the site run: prediction before the rerun
runs/fly_generator_v1 is what the released site shows, and it was recorded against a
screen that has since been rebuilt twice - the relaxation model of E70, then the elemental
anchors and leverage-scaled uncertainty of E72. Its first logged decision is
Mo0.09 V0.91 scored at -85 meV/atom, which is the vanadium-corner false positive E72
traced to the expansion's own 58 meV/atom error at a pure element. The site is currently
showing a generator being rewarded for finding where the model is wrong.
scripts/stage_b.py rewards drive_cold raw, so a rerun without changing it would record
the same artefacts again. The reward becomes the conservative bound drive + sigma, as the
generator campaign already uses.
Predicted, before the rerun:
Vanadium-rich proposals vanish from the accepted set. A composition above about
x_V = 0.85 carries sigma of 110 to 130 meV/atom, so its conservative score is positive
and it cannot be accepted. Falsified if any V-rich composition survives in the top ten.
The best score lands between -110 and -120 meV/atom. The same corrected screen, in
the 58-composition campaign, reached -117 conservative across six seeds.
The family is Mo-Ta-W, optionally with Nb, matching that campaign, and 100 per cent
of the top ten is drawn from {Mo, Nb, Ta, W}.
Distinct-composition count stays near the recorded 55, because the geometry of the
search is unchanged; only what it is paid for moved.
A failure of 1 means the objective did not take effect. A failure of 2 or 3 would mean the
screen rebuild moved the answer rather than only its uncertainty, which would be a finding
in itself.
Outcome. Three of four predictions held; the second is falsified, and the reason is worth
more than the prediction was.
#
predicted
measured
1
no V-rich composition in the top ten
0 of 10
holds
2
best between -110 and -120 meV/atom
-90
falsified
3
top ten drawn only from Mo/Nb/Ta/W
10 of 10
holds
4
distinct count near 55
51.9 +/- 3.7
holds
The objective took effect exactly as intended. The compositions that used to be rewarded are
now penalised rather than merely demoted: Mo0.09 V0.91 moves from -85 to +61 meV/atom and
pure vanadium from -91 to +114. The fly still proposes them at the same rate - 10 of 240 -
because exploration is unchanged; it can no longer be paid for them. AUC_Q is 20.10 +/- 2.00
against 20.34 before, so the harder objective cost the search nothing measurable.
Why the second failed, tested rather than assumed. The predicted range came from the
58-composition campaign, which reached -117. Scoring that campaign's best composition with
the same screen this rerun used gives:
composition
drive
sigma
conservative
Mo0.62 Ta0.38
-151
36
-116
this rerun's best, Mo0.33 Nb0.15 Ta0.21 W0.31
-122
31
-90
The screen agrees with the prediction. The search did not find it.stage_b spends 60
rounds of 4 proposals - 240 - where the campaign spends 800 per seed through a different
move loop, so the site run under-searches by 26 meV/atom against what its own screen rates
reachable. I predicted what was achievable and called it what this configuration would
achieve, which are different claims about two different scripts.
So the site is honest but not impressive, and the gap is budget, not physics. That is
the "long generator run" already sitting in the open items, now with a number attached: 26
meV/atom left on the table at 240 proposals.
The site is rebuilt. The runs index reads 240 screen, best -90 meV where it read
120 screen, best -47 meV/atom. The bundle carries 9 morphology skeletons, 4,114 neurons
and 20,000 edges - the same 9 skeletons as all three exported runs, so the run is not
deficient against the ones it sits beside; showing the whole brain is a change to every run,
not a repair to this one.
Results
EXPERIMENTS.md · line 4305
Outcome. Three of four predictions held; the second is falsified, and the reason is worth
more than the prediction was.
The full record
EXPERIMENTS.md · lines 4276–4344
E79 — Re-recording the site run: prediction before the rerun
runs/fly_generator_v1 is what the released site shows, and it was recorded against a
screen that has since been rebuilt twice - the relaxation model of E70, then the elemental
anchors and leverage-scaled uncertainty of E72. Its first logged decision is
Mo0.09 V0.91 scored at -85 meV/atom, which is the vanadium-corner false positive E72
traced to the expansion's own 58 meV/atom error at a pure element. The site is currently
showing a generator being rewarded for finding where the model is wrong.
scripts/stage_b.py rewards drive_cold raw, so a rerun without changing it would record
the same artefacts again. The reward becomes the conservative bound drive + sigma, as the
generator campaign already uses.
Predicted, before the rerun:
Vanadium-rich proposals vanish from the accepted set. A composition above about
x_V = 0.85 carries sigma of 110 to 130 meV/atom, so its conservative score is positive
and it cannot be accepted. Falsified if any V-rich composition survives in the top ten.
The best score lands between -110 and -120 meV/atom. The same corrected screen, in
the 58-composition campaign, reached -117 conservative across six seeds.
The family is Mo-Ta-W, optionally with Nb, matching that campaign, and 100 per cent
of the top ten is drawn from {Mo, Nb, Ta, W}.
Distinct-composition count stays near the recorded 55, because the geometry of the
search is unchanged; only what it is paid for moved.
A failure of 1 means the objective did not take effect. A failure of 2 or 3 would mean the
screen rebuild moved the answer rather than only its uncertainty, which would be a finding
in itself.
Outcome. Three of four predictions held; the second is falsified, and the reason is worth
more than the prediction was.
#
predicted
measured
1
no V-rich composition in the top ten
0 of 10
holds
2
best between -110 and -120 meV/atom
-90
falsified
3
top ten drawn only from Mo/Nb/Ta/W
10 of 10
holds
4
distinct count near 55
51.9 +/- 3.7
holds
The objective took effect exactly as intended. The compositions that used to be rewarded are
now penalised rather than merely demoted: Mo0.09 V0.91 moves from -85 to +61 meV/atom and
pure vanadium from -91 to +114. The fly still proposes them at the same rate - 10 of 240 -
because exploration is unchanged; it can no longer be paid for them. AUC_Q is 20.10 +/- 2.00
against 20.34 before, so the harder objective cost the search nothing measurable.
Why the second failed, tested rather than assumed. The predicted range came from the
58-composition campaign, which reached -117. Scoring that campaign's best composition with
the same screen this rerun used gives:
composition
drive
sigma
conservative
Mo0.62 Ta0.38
-151
36
-116
this rerun's best, Mo0.33 Nb0.15 Ta0.21 W0.31
-122
31
-90
The screen agrees with the prediction. The search did not find it.stage_b spends 60
rounds of 4 proposals - 240 - where the campaign spends 800 per seed through a different
move loop, so the site run under-searches by 26 meV/atom against what its own screen rates
reachable. I predicted what was achievable and called it what this configuration would
achieve, which are different claims about two different scripts.
So the site is honest but not impressive, and the gap is budget, not physics. That is
the "long generator run" already sitting in the open items, now with a number attached: 26
meV/atom left on the table at 240 proposals.
The site is rebuilt. The runs index reads 240 screen, best -90 meV where it read
120 screen, best -47 meV/atom. The bundle carries 9 morphology skeletons, 4,114 neurons
and 20,000 edges - the same 9 skeletons as all three exported runs, so the run is not
deficient against the ones it sits beside; showing the whole brain is a change to every run,
not a repair to this one.