With its set-up defects fixed, does the diversity-seeking search match the best single-goal search?
Yes. Its best find came within 0.5 meV/atom, it covered 4.8 times more of the map, and it found a second stable family.
In the log: The same head-to-head, with the implementation defects fixed
confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04search baselines0 predictions · 1 result paragraphEXPERIMENTS.md lines 6755–6781, lines 7029–7069
What E115 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E115.svg).
Results
EXPERIMENTS.md · line 7029
E115 result: all four predictions confirmed. E114's verdict is WITHDRAWN.
20,000 screens per arm, three seeds, twelve elements, with the three implementation defects
of E114 corrected:
arm cells coverage best drive best off-corner
cem 21.7 18.1% -95.2 meV +15.7 meV
map-elites 104.3 86.9% -94.7 meV -47.4 meV
Prediction 1 confirmed: CrossEntropy went from -86.0 to -95.2, an improvement of 9.2
against the 20 that would have meant the earlier comparison was budget-starved for both arms.
Prediction 2 confirmed, decisively: the two arms' best finds are 0.5 meV/atom apart,
against the 99 meV gap measured at 1200 evaluations. Prediction 3 confirmed: coverage 86.9
per cent against the 60 threshold, and against 96 per cent reachable. Prediction 4
confirmed: both arms' best is a Ta-Mo binary; no high-VEC composition is reported as a win.
So the 99 meV/atom "price of diversity" in E114 was entirely implementation, not method.
The three defects were guessed descriptor ranges, an archive seeded with 64 draws against 115
reachable cells, and a budget an order of magnitude below what the method needs. Corrected,
MAP-Elites matches a strong CrossEntropy on quality while covering 4.8 times as much of
the behaviour space.
The off-corner column is the finding that matters, and it was not the headline prediction.
Outside the Mo-Nb-Ta-W corner, CrossEntropy's best across three seeds is +15.7 meV/atom -
nothing stable - while MAP-Elites finds -47.4, and finds the same family every time:
A reproducible Mo-Ti-Ta basin, stable by 47 meV/atom, that the scalar optimiser never reaches
in 20,000 evaluations across three independent seeds. That is a 63 meV/atom difference on
exactly the axis the diversity machinery was built for, and it is the first evidence in this
project that the archive finds something the scalar search cannot rather than merely spreading
out.
Caveat, and it is the same one as E118: Mo-Ti-Ta contains titanium, which is prohibited
for liquid-oxygen service. It is a real find for the phase-stability question and is filtered
out by the selection layer, not by the ladder. The two stages disagreeing about it is the
architecture working, not failing.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 6755–6781
E115 — The same head-to-head, with the implementation defects fixed
20,000 screens per arm, three seeds, twelve elements, corrected ranges, 600-draw seeding,
DE variation.
Predicted:
CrossEntropy's best driving force stays near -86 meV/atom. It converged by 1200; a
further 18,800 evaluations should buy it very little. If it improves by more than 20
meV/atom then the earlier comparison was budget-starved for both arms and E114's verdict
was premature.
MAP-Elites' best comes within 30 meV/atom of CrossEntropy's. This is the claim. If
the 99 meV gap was budget and seeding rather than the method, it closes here; if it does
not close, quality-diversity costs real quality on this landscape and that is the finding.
MAP-Elites coverage exceeds 60 per cent against the 96 per cent that is reachable.
Below 40 per cent means the variation operator still cannot move around the space and the
problem is the operator, not the budget.
Neither arm's best find is a high-VEC fcc composition, because the hull says those
are unstable - CrNiCoCu is +72.7 meV/atom. The archive should contain them and score
them as rejections. A run that reports a high-VEC cell as a win is broken, and this
stays the falsification condition: coverage is not success.
Falsified if MAP-Elites still loses by more than 30 meV/atom at twenty thousand
evaluations, in which case the honest recommendation is to keep CrossEntropy as the
optimiser and use the archive only as a reporting structure over what it found - which is
the pending "MAP-Elites archive for multi-basin reporting" task, and a much weaker claim than
replacing the search.
EXPERIMENTS.md · lines 7029–7069
E115 result: all four predictions confirmed. E114's verdict is WITHDRAWN.
20,000 screens per arm, three seeds, twelve elements, with the three implementation defects
of E114 corrected:
arm cells coverage best drive best off-corner
cem 21.7 18.1% -95.2 meV +15.7 meV
map-elites 104.3 86.9% -94.7 meV -47.4 meV
Prediction 1 confirmed: CrossEntropy went from -86.0 to -95.2, an improvement of 9.2
against the 20 that would have meant the earlier comparison was budget-starved for both arms.
Prediction 2 confirmed, decisively: the two arms' best finds are 0.5 meV/atom apart,
against the 99 meV gap measured at 1200 evaluations. Prediction 3 confirmed: coverage 86.9
per cent against the 60 threshold, and against 96 per cent reachable. Prediction 4
confirmed: both arms' best is a Ta-Mo binary; no high-VEC composition is reported as a win.
So the 99 meV/atom "price of diversity" in E114 was entirely implementation, not method.
The three defects were guessed descriptor ranges, an archive seeded with 64 draws against 115
reachable cells, and a budget an order of magnitude below what the method needs. Corrected,
MAP-Elites matches a strong CrossEntropy on quality while covering 4.8 times as much of
the behaviour space.
The off-corner column is the finding that matters, and it was not the headline prediction.
Outside the Mo-Nb-Ta-W corner, CrossEntropy's best across three seeds is +15.7 meV/atom -
nothing stable - while MAP-Elites finds -47.4, and finds the same family every time:
A reproducible Mo-Ti-Ta basin, stable by 47 meV/atom, that the scalar optimiser never reaches
in 20,000 evaluations across three independent seeds. That is a 63 meV/atom difference on
exactly the axis the diversity machinery was built for, and it is the first evidence in this
project that the archive finds something the scalar search cannot rather than merely spreading
out.
Caveat, and it is the same one as E118: Mo-Ti-Ta contains titanium, which is prohibited
for liquid-oxygen service. It is a real find for the phase-stability question and is filtered
out by the selection layer, not by the ladder. The two stages disagreeing about it is the
architecture working, not failing.
Related entries
E114 — The search space was eight elements, and the search collapses inside it
E118 — Multi-objective selection over the feasible set