Experiments · E79

Once the model's vanadium error stops paying, what does the site's generator run find?

Partly. Vanadium-rich picks vanished and the family stayed Mo–Nb–Ta–W, but the best score was −90 meV/atom, short of the predicted −110 to −120.

In the log: Re-recording the site run: prediction before the rerun

mixedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04generator · fly brain0 predictions · 1 result paragraphEXPERIMENTS.md lines 4276–4344
exp E79 diagram
What E79 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E79.svg).

Pre-registration

The pre-registration, as written

E79 — Re-recording the site run: prediction before the rerun

runs/fly_generator_v1 is what the released site shows, and it was recorded against a screen that has since been rebuilt twice - the relaxation model of E70, then the elemental anchors and leverage-scaled uncertainty of E72. Its first logged decision is Mo0.09 V0.91 scored at -85 meV/atom, which is the vanadium-corner false positive E72 traced to the expansion's own 58 meV/atom error at a pure element. The site is currently showing a generator being rewarded for finding where the model is wrong.

scripts/stage_b.py rewards drive_cold raw, so a rerun without changing it would record the same artefacts again. The reward becomes the conservative bound drive + sigma, as the generator campaign already uses.

Predicted, before the rerun:

  1. Vanadium-rich proposals vanish from the accepted set. A composition above about x_V = 0.85 carries sigma of 110 to 130 meV/atom, so its conservative score is positive and it cannot be accepted. Falsified if any V-rich composition survives in the top ten.
  2. The best score lands between -110 and -120 meV/atom. The same corrected screen, in the 58-composition campaign, reached -117 conservative across six seeds.
  3. The family is Mo-Ta-W, optionally with Nb, matching that campaign, and 100 per cent of the top ten is drawn from {Mo, Nb, Ta, W}.
  4. Distinct-composition count stays near the recorded 55, because the geometry of the search is unchanged; only what it is paid for moved.

A failure of 1 means the objective did not take effect. A failure of 2 or 3 would mean the screen rebuild moved the answer rather than only its uncertainty, which would be a finding in itself.

Outcome. Three of four predictions held; the second is falsified, and the reason is worth more than the prediction was.

# predicted measured
1 no V-rich composition in the top ten 0 of 10 holds
2 best between -110 and -120 meV/atom -90 falsified
3 top ten drawn only from Mo/Nb/Ta/W 10 of 10 holds
4 distinct count near 55 51.9 +/- 3.7 holds

The objective took effect exactly as intended. The compositions that used to be rewarded are now penalised rather than merely demoted: Mo0.09 V0.91 moves from -85 to +61 meV/atom and pure vanadium from -91 to +114. The fly still proposes them at the same rate - 10 of 240 - because exploration is unchanged; it can no longer be paid for them. AUC_Q is 20.10 +/- 2.00 against 20.34 before, so the harder objective cost the search nothing measurable.

Why the second failed, tested rather than assumed. The predicted range came from the 58-composition campaign, which reached -117. Scoring that campaign's best composition with the same screen this rerun used gives:

composition drive sigma conservative
Mo0.62 Ta0.38 -151 36 -116
this rerun's best, Mo0.33 Nb0.15 Ta0.21 W0.31 -122 31 -90

The screen agrees with the prediction. The search did not find it. stage_b spends 60 rounds of 4 proposals - 240 - where the campaign spends 800 per seed through a different move loop, so the site run under-searches by 26 meV/atom against what its own screen rates reachable. I predicted what was achievable and called it what this configuration would achieve, which are different claims about two different scripts.

So the site is honest but not impressive, and the gap is budget, not physics. That is the "long generator run" already sitting in the open items, now with a number attached: 26 meV/atom left on the table at 240 proposals.

The site is rebuilt. The runs index reads 240 screen, best -90 meV where it read 120 screen, best -47 meV/atom. The bundle carries 9 morphology skeletons, 4,114 neurons and 20,000 edges - the same 9 skeletons as all three exported runs, so the run is not deficient against the ones it sits beside; showing the whole brain is a change to every run, not a repair to this one.

Results

EXPERIMENTS.md · line 4305

Outcome. Three of four predictions held; the second is falsified, and the reason is worth more than the prediction was.

The full record

EXPERIMENTS.md · lines 4276–4344

E79 — Re-recording the site run: prediction before the rerun

runs/fly_generator_v1 is what the released site shows, and it was recorded against a screen that has since been rebuilt twice - the relaxation model of E70, then the elemental anchors and leverage-scaled uncertainty of E72. Its first logged decision is Mo0.09 V0.91 scored at -85 meV/atom, which is the vanadium-corner false positive E72 traced to the expansion's own 58 meV/atom error at a pure element. The site is currently showing a generator being rewarded for finding where the model is wrong.

scripts/stage_b.py rewards drive_cold raw, so a rerun without changing it would record the same artefacts again. The reward becomes the conservative bound drive + sigma, as the generator campaign already uses.

Predicted, before the rerun:

  1. Vanadium-rich proposals vanish from the accepted set. A composition above about x_V = 0.85 carries sigma of 110 to 130 meV/atom, so its conservative score is positive and it cannot be accepted. Falsified if any V-rich composition survives in the top ten.
  2. The best score lands between -110 and -120 meV/atom. The same corrected screen, in the 58-composition campaign, reached -117 conservative across six seeds.
  3. The family is Mo-Ta-W, optionally with Nb, matching that campaign, and 100 per cent of the top ten is drawn from {Mo, Nb, Ta, W}.
  4. Distinct-composition count stays near the recorded 55, because the geometry of the search is unchanged; only what it is paid for moved.

A failure of 1 means the objective did not take effect. A failure of 2 or 3 would mean the screen rebuild moved the answer rather than only its uncertainty, which would be a finding in itself.

Outcome. Three of four predictions held; the second is falsified, and the reason is worth more than the prediction was.

# predicted measured
1 no V-rich composition in the top ten 0 of 10 holds
2 best between -110 and -120 meV/atom -90 falsified
3 top ten drawn only from Mo/Nb/Ta/W 10 of 10 holds
4 distinct count near 55 51.9 +/- 3.7 holds

The objective took effect exactly as intended. The compositions that used to be rewarded are now penalised rather than merely demoted: Mo0.09 V0.91 moves from -85 to +61 meV/atom and pure vanadium from -91 to +114. The fly still proposes them at the same rate - 10 of 240 - because exploration is unchanged; it can no longer be paid for them. AUC_Q is 20.10 +/- 2.00 against 20.34 before, so the harder objective cost the search nothing measurable.

Why the second failed, tested rather than assumed. The predicted range came from the 58-composition campaign, which reached -117. Scoring that campaign's best composition with the same screen this rerun used gives:

composition drive sigma conservative
Mo0.62 Ta0.38 -151 36 -116
this rerun's best, Mo0.33 Nb0.15 Ta0.21 W0.31 -122 31 -90

The screen agrees with the prediction. The search did not find it. stage_b spends 60 rounds of 4 proposals - 240 - where the campaign spends 800 per seed through a different move loop, so the site run under-searches by 26 meV/atom against what its own screen rates reachable. I predicted what was achievable and called it what this configuration would achieve, which are different claims about two different scripts.

So the site is honest but not impressive, and the gap is budget, not physics. That is the "long generator run" already sitting in the open items, now with a number attached: 26 meV/atom left on the table at 240 proposals.

The site is rebuilt. The runs index reads 240 screen, best -90 meV where it read 120 screen, best -47 meV/atom. The bundle carries 9 morphology skeletons, 4,114 neurons and 20,000 edges - the same 9 skeletons as all three exported runs, so the run is not deficient against the ones it sits beside; showing the whole brain is a change to every run, not a repair to this one.

Related entries

Built with PRISMWebsite and visualizations made using Claude