Experiments · E63

Can the generators be compared on the full requirement instead of the quick check?

No. On equilibrium alone nearly every composition orders inside 90–1000 K, so every generator scores zero or near zero; kinetics must decide.

In the log: Gating on the equilibrium terms sends every generator to zero, and that is the specification talking

recordedDate not stated in the log; it was written between the commit of 2026-09-13 08:16 and the first commit that contains it, 2026-09-16 02:04unclassified0 predictions · 0 result paragraphsEXPERIMENTS.md lines 3430–3476
exp E63 diagram
What E63 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E63.svg).

Results

No result paragraph for this entry was found in the log.

The full record

EXPERIMENTS.md · lines 3430–3476

E63 — Gating on the equilibrium terms sends every generator to zero, and that is the specification talking

E60 scored the generators on the free rung alone, which answers one of the three things that can break the requirement. The obvious repair is to keep the cheap reward and let the dear rungs decide what counts as a find - trial.run takes a confirm argument for exactly this. Gating on the two equilibrium terms, off-lattice stability and no ordering inside the window:

arm screen only gated on equilibrium
elitist 4.63 0.00
cem 3.24 0.00
fly 6.29 0.24 +/- 0.24
fly-frozen 1.61 0.24 +/- 0.24
sparse 0.65 0.00
uniform 0.00 0.00

paired fly minus frozen +0.00, one seed of four. No arm can be told from another.

This is not the arms failing. Of 348 compositions measured - 343 across the sampled box and all five later examined outside it, up to Ta 0.44 - every one has an ordering transition inside 90 to 1000 K. The gate is the equilibrium criterion and the equilibrium criterion is empty, so it rejects everything by construction, including the alloy with a five-week experiment behind it (E58).

The harness said so in advance. test_the_harness_is_blind_when_the_target_is_too_narrow asserts that where qualifiers are isolated every arm scores the same and a comparison measures luck rather than learning. Gated this way, the real problem is that regime.

What follows for the design. The kinetic term is not a refinement of the verdict, it is the verdict; and it costs twenty minutes a composition. That leaves three options and no free one:

  • score on the screen and accept that the comparison covers one term of three (E61, which is where the learning result comes from);
  • gate on all three and pay twenty minutes per candidate, which no arm-versus-arm comparison in this project can currently afford;
  • build a cheap kinetic proxy. A rule of mixtures over pure-element migration barriers reaches r = +0.77 against measured E_m on five compositions, with a residual of 0.25 eV - a factor of four in diffusion distance at 1000 K, and five points cannot validate anything. Settling it needs the activation energy computed across the Mo-Nb-Ta-W corner.

The working pattern meanwhile is the ladder as designed and as used by hand in E60 and E62: the generator optimises the screen, the few compositions it likes best are carried up, and the survivors are reported with the rung that confirmed them. That produced a candidate that survives all three. It is not the same thing as a generator comparison on the full verdict, and this entry exists so the two are not confused later.

Related entries

Built with PRISMWebsite and visualizations made using Claude