Experiments · E115

With its set-up defects fixed, does the diversity-seeking search match the best single-goal search?

Yes. Its best find came within 0.5 meV/atom, it covered 4.8 times more of the map, and it found a second stable family.

In the log: The same head-to-head, with the implementation defects fixed

confirmedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04search baselines0 predictions · 1 result paragraphEXPERIMENTS.md lines 6755–6781, lines 7029–7069
exp E115 diagram
What E115 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E115.svg).

Results

EXPERIMENTS.md · line 7029

E115 result: all four predictions confirmed. E114's verdict is WITHDRAWN.

20,000 screens per arm, three seeds, twelve elements, with the three implementation defects of E114 corrected:

arm           cells   coverage   best drive   best off-corner
cem            21.7      18.1%    -95.2 meV        +15.7 meV
map-elites    104.3      86.9%    -94.7 meV        -47.4 meV

Prediction 1 confirmed: CrossEntropy went from -86.0 to -95.2, an improvement of 9.2 against the 20 that would have meant the earlier comparison was budget-starved for both arms. Prediction 2 confirmed, decisively: the two arms' best finds are 0.5 meV/atom apart, against the 99 meV gap measured at 1200 evaluations. Prediction 3 confirmed: coverage 86.9 per cent against the 60 threshold, and against 96 per cent reachable. Prediction 4 confirmed: both arms' best is a Ta-Mo binary; no high-VEC composition is reported as a win.

So the 99 meV/atom "price of diversity" in E114 was entirely implementation, not method. The three defects were guessed descriptor ranges, an archive seeded with 64 draws against 115 reachable cells, and a budget an order of magnitude below what the method needs. Corrected, MAP-Elites matches a strong CrossEntropy on quality while covering 4.8 times as much of the behaviour space.

The off-corner column is the finding that matters, and it was not the headline prediction. Outside the Mo-Nb-Ta-W corner, CrossEntropy's best across three seeds is +15.7 meV/atom - nothing stable - while MAP-Elites finds -47.4, and finds the same family every time:

map-elites  Mo0.52 Ti0.25 Ta0.23   -47.1
map-elites  Mo0.63 Ti0.26 Ta0.11   -46.5
map-elites  Mo0.56 Ti0.27 Ta0.14   -48.7
cem         W0.31 Ta0.26 Mo0.12 V0.11 Cr0.10 Nb0.06   +29.6

A reproducible Mo-Ti-Ta basin, stable by 47 meV/atom, that the scalar optimiser never reaches in 20,000 evaluations across three independent seeds. That is a 63 meV/atom difference on exactly the axis the diversity machinery was built for, and it is the first evidence in this project that the archive finds something the scalar search cannot rather than merely spreading out.

Caveat, and it is the same one as E118: Mo-Ti-Ta contains titanium, which is prohibited for liquid-oxygen service. It is a real find for the phase-stability question and is filtered out by the selection layer, not by the ladder. The two stages disagreeing about it is the architecture working, not failing.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 6755–6781

E115 — The same head-to-head, with the implementation defects fixed

20,000 screens per arm, three seeds, twelve elements, corrected ranges, 600-draw seeding, DE variation.

Predicted:

  1. CrossEntropy's best driving force stays near -86 meV/atom. It converged by 1200; a further 18,800 evaluations should buy it very little. If it improves by more than 20 meV/atom then the earlier comparison was budget-starved for both arms and E114's verdict was premature.
  2. MAP-Elites' best comes within 30 meV/atom of CrossEntropy's. This is the claim. If the 99 meV gap was budget and seeding rather than the method, it closes here; if it does not close, quality-diversity costs real quality on this landscape and that is the finding.
  3. MAP-Elites coverage exceeds 60 per cent against the 96 per cent that is reachable. Below 40 per cent means the variation operator still cannot move around the space and the problem is the operator, not the budget.
  4. Neither arm's best find is a high-VEC fcc composition, because the hull says those are unstable - CrNiCoCu is +72.7 meV/atom. The archive should contain them and score them as rejections. A run that reports a high-VEC cell as a win is broken, and this stays the falsification condition: coverage is not success.

Falsified if MAP-Elites still loses by more than 30 meV/atom at twenty thousand evaluations, in which case the honest recommendation is to keep CrossEntropy as the optimiser and use the archive only as a reporting structure over what it found - which is the pending "MAP-Elites archive for multi-basin reporting" task, and a much weaker claim than replacing the search.

EXPERIMENTS.md · lines 7029–7069

E115 result: all four predictions confirmed. E114's verdict is WITHDRAWN.

20,000 screens per arm, three seeds, twelve elements, with the three implementation defects of E114 corrected:

arm           cells   coverage   best drive   best off-corner
cem            21.7      18.1%    -95.2 meV        +15.7 meV
map-elites    104.3      86.9%    -94.7 meV        -47.4 meV

Prediction 1 confirmed: CrossEntropy went from -86.0 to -95.2, an improvement of 9.2 against the 20 that would have meant the earlier comparison was budget-starved for both arms. Prediction 2 confirmed, decisively: the two arms' best finds are 0.5 meV/atom apart, against the 99 meV gap measured at 1200 evaluations. Prediction 3 confirmed: coverage 86.9 per cent against the 60 threshold, and against 96 per cent reachable. Prediction 4 confirmed: both arms' best is a Ta-Mo binary; no high-VEC composition is reported as a win.

So the 99 meV/atom "price of diversity" in E114 was entirely implementation, not method. The three defects were guessed descriptor ranges, an archive seeded with 64 draws against 115 reachable cells, and a budget an order of magnitude below what the method needs. Corrected, MAP-Elites matches a strong CrossEntropy on quality while covering 4.8 times as much of the behaviour space.

The off-corner column is the finding that matters, and it was not the headline prediction. Outside the Mo-Nb-Ta-W corner, CrossEntropy's best across three seeds is +15.7 meV/atom - nothing stable - while MAP-Elites finds -47.4, and finds the same family every time:

map-elites  Mo0.52 Ti0.25 Ta0.23   -47.1
map-elites  Mo0.63 Ti0.26 Ta0.11   -46.5
map-elites  Mo0.56 Ti0.27 Ta0.14   -48.7
cem         W0.31 Ta0.26 Mo0.12 V0.11 Cr0.10 Nb0.06   +29.6

A reproducible Mo-Ti-Ta basin, stable by 47 meV/atom, that the scalar optimiser never reaches in 20,000 evaluations across three independent seeds. That is a 63 meV/atom difference on exactly the axis the diversity machinery was built for, and it is the first evidence in this project that the archive finds something the scalar search cannot rather than merely spreading out.

Caveat, and it is the same one as E118: Mo-Ti-Ta contains titanium, which is prohibited for liquid-oxygen service. It is a real find for the phase-stability question and is filtered out by the selection layer, not by the ladder. The two stages disagreeing about it is the architecture working, not failing.

Related entries

Built with PRISMWebsite and visualizations made using Claude