Does the quick check's relaxation estimate work where the generator actually searches?
No. There it predicts about 10–12 meV/atom everywhere while the truth varies by 38; it is off by −24 on average.
In the log: The screen's relaxation model does no work where the search actually is
mixedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04unclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 3744–3781
What E69 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E69.svg).
Results
No result paragraph for this entry was found in the log.
The full record
EXPERIMENTS.md · lines 3744–3781
E69 — The screen's relaxation model does no work where the search actually is
The free rung estimates the relaxation an expansion cannot represent, from the square of the
atomic size misfit, with a leave-one-out error of 18 meV/atom across the design space (E59).
Thirteen compositions have since been relaxed properly, which makes that claim checkable
out of sample in the region the generator converged on.
Predicted: the bias would grow with distance from the fitted box, so it should track the
largest element fraction, and correcting for it should bring the error under 10 meV.
value
bias, mean
-24 meV/atom, one-signed
bias, spread
13 meV/atom
correlation with largest fraction
-0.58
correlation with size misfit
+0.05
correlation with element count
+0.49
The direction holds and the mechanism does not. Size misfit spans 2.32 to 2.52 per cent
across all thirteen, so the model predicts 9.9 to 11.6 meV of relaxation for every one of
them - a spread of 1.7 - while the truth moves by 38. The estimate is not wrong so much
as blind: it was fitted across a design space where misfit runs from 2.5 to 7 per cent, and
inside the Mo-Nb-Ta-W corner that variable is effectively constant.
Two nearly identical compositions make the point: Ta.47 Mo.41 Nb.12 is under-reported by 8
meV/atom and Ta.46 Mo.37 Nb.18 by 37. The measurement scatter is 2 meV on the relaxation and
1 on the screen, so that 29 meV gap is real.
And it is not predictable. A leave-one-out fit on the largest fraction and the element
count moves the residual from 13.3 to 12.4 meV. A constant offset changes nothing at all for
ranking, which is the only thing the screen is asked to do.
What follows. The screen's honest description in this region is a one-signed offset of
24 meV and a spread of 13 - and one-signed toward under-reporting, so it never promises a
stability that is not there. Thirteen meV against a forty meV qualifying threshold still
ranks, and doing better means relaxing, which is the next rung's job. The docstring now says
this instead of quoting the 18 meV figure, which is true of the design space and misleading
about the search.
Related entries
E59 — The question that decides costs four milliseconds, and twenty thousand compositions say…