Experiments · E197

Does searching on the refitted energy model give deeper, more varied finds?

No. The finds' spread fell to about 15 meV/atom against a bar of 40, and the best find got shallower, not deeper.

In the log: generate on v6b (2026-09-19 02:43; queued behind the v6 pair)

mixedDate 2026-09-19 02:43, as written in the logrung 1 · ordering0 predictions · 0 result paragraphsEXPERIMENTS.md lines 12636–12646, lines 12978–12984
exp E197 diagram
What E197 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E197.svg).

Pre-registration

The pre-registration, as written

E197 — generate on v6b. elitist 147.1 / 261.5 finds / depth −64.2 / best −89.1; fly 109.3 / 214 / −56.7 / −88.1. Bars: spread ≥ 40 → 15.9 and 13.6, falsified; best ≤ −113 → −88, falsified; elitist ≥ fly on rate → confirmed. v6b's landscape is shallower and narrower than v5's by ~15 meV at the bottom — consistent with E190b (the refit raised the ordered states) — and the fly loses more from it than elitist does (214 vs 284 finds; elitist 261 vs 246). The fly is more sensitive to the reward model than the model-free arm. v6b is not the deployed model; v5 stays.

Results

No result paragraph for this entry was found in the log.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 12636–12646

E197 — generate on v6b (2026-09-19 02:43; queued behind the v6 pair)

N5 as written, with two changes forced by tonight: the reward model is v6b (the deep-weighted refit; v6a inherits v5's regression), and the arm that matters alongside the fly is elitist, since E195b/E198 make it the arm to beat. Same settings as E195b, eight-element space, 2 seeds each. Predictions: 1. the finds' v6b energy spread stays ≥ 40 meV and the best find is ≥ 10 meV below E195b's best (−103): the loop closed once and the reward moved. 2. elitist ≥ fly on AUC_Q again — a refit does not change which arm searches better. 3. v6b's E190b MoNbTaVW T_c ≥ 500 K (E192's prediction 3) is a precondition; if v6b fails it, E197 runs anyway and is read as "generate on a v6 that did not fix rung 1". Chain runs/e197_chain.sh.

EXPERIMENTS.md · lines 12978–12984

E197 — generate on v6b. elitist 147.1 / 261.5 finds / depth −64.2 / best −89.1; fly 109.3 / 214 / −56.7 / −88.1. Bars: spread ≥ 40 → 15.9 and 13.6, falsified; best ≤ −113 → −88, falsified; elitist ≥ fly on rate → confirmed. v6b's landscape is shallower and narrower than v5's by ~15 meV at the bottom — consistent with E190b (the refit raised the ordered states) — and the fly loses more from it than elitist does (214 vs 284 finds; elitist 261 vs 246). The fly is more sensitive to the reward model than the model-free arm. v6b is not the deployed model; v5 stays.

Related entries

Built with PRISMWebsite and visualizations made using Claude