The fidelity ladder

The ladder

The generator, this project’s search program, proposes alloys. Each one is asked a single question: does it stay one bcc solid solution everywhere from 90 to 1000 K? The ladder asks it five times, with different physics at each rung. Most rungs cost more than the one below, and most candidates stop early.

In a random solid solution the elements share one crystal lattice, scattered at random over its sites. Here the lattice is bcc (body-centred cubic): cubes with an atom at each corner and one at the centre. The candidates are compositions over the nine elements the energy model covers: Cr, Hf, Mo, Nb, Ta, Ti, V, W and Zr.

Each rung below opens with a plain question and a short answer. The technical detail follows, then the rung’s known answer and what it is blind to. A known answer is a result known in advance, which the rung had to reproduce before it was trusted. Three recorded runs are then replayed: an alloy cooling on rung 1, the choice of the rung-0 model, and the search’s 32 finds walked up the ladder.

ladder diagram
The fidelity ladder, drawn with MetaPost and the fiziko library. Tap to enlarge.

Five rungs

The chart shows what each rung costs per composition, as measured. Each rung’s cost, error and blind spots come from forager/ladder/spec.py, which records only what has been measured.

Rung 0 · Screen0.1–0.5 s
Rung 1 · Ordering15–69 min
Rung 2 · Lattice8–34 s
Rung 3 · Kinetics52 min · one measurement
Rung 4 · First principles1.2–1.4 h
0.1 s1 s10 s1 min10 min1 h10 h
Time per composition, read from forager/ladder/spec.py. The axis is logarithmic: equal distances mean equal ratios of time. Each bar spans the measured range that the rung’s cost_source states. The dark tick is its single typical figure (cost_seconds). Rung 1 was measured on E240’s 32 finds, rung 4 on a rented A100 GPU. Rung 2 costs less than rung 1.

Rung 0 · 0.1–0.5 s · typical 0.3 s run inside the search program, with the model loaded once

Screen

How much lower in energy is the mixed alloy than its alternatives?

Rung 0 answers in a fraction of a second. A fast energy model estimates the energy of the alloy with its atoms mixed at random on the bcc lattice. It compares that with the same atoms as pure elements, and with the cheapest mixture of competing phases.

Precisely, rung 0 gives the formation energy of the random bcc solid solution at this composition. Formation energy is the alloy’s energy minus that of the same atoms as separate pure elements. Negative means the atoms prefer to mix. The pure elements are taken as bcc, on the energy scale (the frame) of the DFT training data. Rung 0 also gives how far the alloy lies below the off-lattice hull, the cheapest mixture of competing phases, explained under rung 2.

The model is an embedded cluster expansion (pyeCE), fitted to DFT energies from the RHEA database. DFT, density functional theory, is the quantum-mechanical calculation of rung 4. A cluster expansion writes the energy of an arrangement as a sum over small groups of sites (diagram below). “Embedded” means each element is described by a short list of learned numbers: six of them in “embedding 6”. The DFT energies a model is fitted to are called its labels.

The model that ships is E237’s embedding-6 ensemble, fitted on the original labels. The ensemble is three fits of one model that differ only in their random starting point (the seed). Rung 0 uses their average; where the three disagree, the model is unsure.

E238 refitted the model on corrected labels, in several versions (arms). Every arm over-bound the search’s random cells, the three DFT cells from E223’s finds described below. Over-binding means putting them lower in energy than DFT does. So E238’s pre-registered fallback applies, the plan written down before the run: keep the original model. The generator’s reward, the score that steers its search, comes from this rung alone.

cluster expansion diagram
The idea rung 0 builds on, a cluster expansion. The energy of an arrangement is written as a weighted sum over small groups of sites: single sites, pairs, triangles.
ece diagram
Rung 0’s model, left to right: the clusters around each atom, then each element turned into six learned numbers. These are combined by cluster shape and fed to a small neural network (a flexible fitted function). The last step is the average over all cells.
labels diagram
Where its training energies (labels) come from. A DFT snapshot whose atoms have moved is mapped back onto ideal lattice sites. The energy of that move is corrected for.

Known answerEleven independent QE cells, which no model chose. QE is Quantum ESPRESSO, the DFT program of rung 4. The cells are B2 Mo–Ta and three random Mo–Ta cells; an ordered and three random MoNbTaVW cells (E191); and random cells from E222, E217 and E183. B2 is the ordered pattern in which every nearest neighbour is of the other element. On these cells the ensemble’s mean absolute error, the average size of its miss, is 7.5 meV/atom. (A meV is a thousandth of an electronvolt.) v5, the embedding-3 model it replaced, scored 20.6. On the three random cells of E223’s own finds, computed in DFT, the mean error is 6.1 (E237).

Blind toMagnetism. The model ignores electron spin. Yet chromium turns antiferromagnetic (neighbouring magnetic moments point opposite ways) below its Néel point, 311 K, which is inside the window. Ignoring spin puts antiferromagnetic Cr off by 12.4 meV/atom (E216). Every compound of three or more elements: the hull’s 671 phases stop at binaries, compounds of two elements (E147). Relaxation: atoms shifting off their ideal lattice sites. The pure-element references: the energies of pure Ti, Cr and Hf in the training data (the label frame) are not yet pinned down. The error this leaves grows in step with each element’s share (it is linear in composition). So it moves absolute formation energies but cancels at fixed composition, which is what rung 1 uses. E239 is measuring them.

Rung 1 · 15–69 min · typical 32 min on one processor core, a 432-site cell, measured over E240’s 32 finds

Ordering

Does the alloy’s arrangement change between 90 and 1000 K?

Rung 1 slowly cools a computer model of the alloy and watches for the moment its atoms sort themselves into a regular pattern. If that ordering happens inside the window, the alloy fails.

Precisely: at what temperature does this composition order, on the physical temperature axis? “Physical” means the true temperature; one sampler on this page reports half of it (see “Rung 1, watched”).

The method is Metropolis Monte Carlo: a simulated slow cooling made of random trial swaps of two atoms (diagram below). The swaps are canonical: they never change the composition. The cell is a 6×6×6 block of bcc cubes, 432 sites. The run:

  • one cooling leg, from 2600 K down to 100 K, then on to 80 and 60 K;
  • 100 + 200 sweeps per temperature (100 to settle, 200 to measure); one sweep is 432 attempted swaps.

The energy is the average of the three rung-0 fits, each computed exactly by a faster re-implementation (ExactECE). So rung 0 and rung 1 score one model.

The transition shows as a peak in the heat capacity, read in two ways (channels). The estimate is the peak of dE/dT, the slope of energy against temperature. When the slope finds no peak, the energy-fluctuation channel is read instead: how much the energy fluctuates at each temperature.

A transition inside 90–1000 K fails. A composition with no transition down to 60 K is scored as bounded at that floor. The temperature step above the floor is its uncertainty.

metropolis diagram
Rung 1’s one move: pick two atoms and swap them. Keep the swap with probability min(1, e−ΔE/kT), where ΔE is the energy change, T the temperature and k Boltzmann’s constant. So a swap that lowers the energy is always kept; one that raises it is kept more rarely the colder it is. Repeat at each temperature.
sro diagram
How order is read, with the Warren–Cowley parameter α: look at an atom’s eight nearest neighbours. Mostly the other kind gives α below 0 (ordering); what chance alone gives is α = 0; mostly the same kind is above 0.
odt diagram
What rung 1 looks for: the temperature where the lattice goes from random above to ordered below. It is marked by a peak in the heat capacity, the heat it takes to warm the alloy by one kelvin.

Known answerRung 1 was checked against these answers, known in advance:

  • An exact nearest-neighbour Ising model, the textbook model of ordering on a lattice. For an infinite crystal (the thermodynamic limit) it orders at 1410 K. On this cell size the sampler peaks at 1294–1340 K: a finite-size offset, caused by the cell being finite (E215c).
  • Equiatomic Mo–Ta (equal parts) on the shipped model. The bar, set in advance (pre-registered), was a transition inside 500–2600 K. Rung 1 gives 1149 ± 114 K, into perfect B2 (α1 = −1.00). A published cluster expansion with Monte Carlo (CE + MC) gives 2020 ± 545 K (E242 (4)).
  • The exact evaluator reproduces pyeCE’s energy per cell to 6.3 × 10−7 eV on the production cell. It is 54× faster at embedding 9 (E237).
  • Published ordering temperatures (the scorecard). On v5, six of six systems with a comparable published temperature came within 22 % (E210b).
  • The outlier, Cr–Ta–Ti–W, came out at 2.92× its published temperature. DFT shows the model has the right ordering pair, but 2.4× too much ordering energy (E222).

Six conventions the numbers rest on are proved in Lean 4, a program that checks every step of a mathematical proof. Each is tied to tests of the production functions (formal/). The six:

  • the factor of two in pyeCE’s temperature, explained under “Rung 1, watched”;
  • detailed balance of the sweep: at equilibrium, swaps from one arrangement to another balance the swaps back;
  • the decoding of pyeCE’s short-range-order (SRO) output;
  • the Warren–Cowley bound, the lowest value α can take, reached only by perfect order;
  • the labeller’s volume miss: a missed volume minimum can only raise a label;
  • the exact evaluator’s feature fold, which folds the network’s first layer into the cluster functions.

Blind toEverything rung 0 is blind to. Ordered patterns too large for a 6×6×6 cell (superstructures): Mo–Ta’s true ground state, its lowest-energy arrangement, is oC12, 1 meV/atom below B2. Off-lattice order: every arrangement stays on the ideal bcc sites. Transitions above 2600 K, where the run starts: they are cut off (censored) and reported as such. Hysteresis, a transition that differs between heating and cooling: the run has one leg only, cooling. False peaks at the cold end: both channels have artefacts where the sampler freezes and swaps are almost never accepted (E242 (4)). An element below about 7 % of the cell cannot show detectable order, even when perfectly ordered (E240).

Rung 2 · 8–34 s · typical 17 s MACE-MPA-0 on an ordinary processor (CPU), 54-atom cells

Lattice

Is it bcc at all once the atoms can move?

Rung 2 lets the atoms shift off their ideal sites, using MACE, a fast energy model fitted to quantum-mechanical calculations. It then asks whether a mixture of other crystal structures would be lower in energy. If so, the alloy would rather split into them.

MACE is a machine-learned interatomic potential: it gives each atom an energy from the positions of its neighbours (diagram below). Rung 2 relaxes a bcc cell in MACE, letting the atoms and the cell shape settle into the nearest energy minimum. It reports that bcc energy and its spread over random starting arrangements (seeds).

It also reports the driving force against the off-lattice hull at both ends of the window. The hull is the lowest energy the same atoms could reach as a mixture of other phases, including ones that are not bcc. The driving force is the alloy’s height above that floor. A positive drive at the cold end means the solution sits above the cheapest mixture of other phases.

hull diagram
Rung 2’s question. The floor is the convex hull, the lowest energy reachable as a mixture of competing phases. A candidate above the floor gives up energy by splitting into those phases; its height above the floor is the driving force.
laves diagram
Why rung 2 exists: at HfV2 a Laves phase, with separate sites for big and small atoms, sits 123 meV/atom below the bcc solution. A model that knows only the bcc lattice cannot see it.
mace diagram
How MACE scores a structure: each atom’s energy comes from the neighbours inside a cutoff distance, through their distances and angles. The total is the sum over atoms.

Known answerMACE and Quantum ESPRESSO agree to 1 and 4 meV/atom on the gaps in mixing energy between alloys. Those gaps are what the ranking uses. For comparison, DFT’s own k-point error is about 4 (E78); k-points are the grid on which DFT samples the electrons’ waves. MACE’s absolute energies carry a 28 meV/atom offset. The offset is constant, so it cancels in the gaps.

Blind toConfigurational entropy, the stability a random alloy gains from its many possible arrangements: one relaxed cell is one arrangement. Magnetism. The hull it compares against holds no compound beyond binaries (two elements), so a multi-element composition can look more stable than it is.

Rung 3 · 52 min · one measurement on MoNbTaW: 24 vacancy sites, 30 hop paths (NEB bands)

Kinetics

Can the atoms move far enough, in the part’s working life, for any change to happen?

Atoms in a metal move by hopping into empty sites (vacancies). Rung 3 measures how hard those hops are, and how far atoms can travel in the service life. If even the fastest cannot travel far enough, the alloy cannot change, whatever its energies prefer. For now the rung records its numbers without passing or failing anything.

Precisely: over the service life, is the fastest transport this composition can muster still too slow for the transformation the thermodynamics wants?

Rung 3 computes the spectrum of vacancy-migration barriers: the spread of energy humps an atom climbs to hop into a vacancy. Each barrier comes from a nudged-elastic-band (NEB) calculation, which finds the easiest path for one hop; 30 converged bands are used. It also computes the vacancy formation energy for each element (species): the energy to remove one of its atoms.

From these it estimates how far an atom travels in the service life, as a range (an interval) rather than one number. The pass-or-fail check (the gate) reads the most mobile end of that range.

diffusion diagram
Rung 3’s question: an atom moves only by hopping into a vacancy, an empty site. The barrier to that hop decides whether order or decomposition (splitting into other phases) can happen within the service life.

Known answerThe rung’s first version gave a single number (a point estimate), and a test proved it wrong (E92). It was 1.58 eV off, in the unsafe direction: it made alloys look more frozen than they are. The rung was rebuilt. On pure Mo its vacancy formation energy agrees with DFT: 3.107 eV by MACE, 3.164 by DFT (E219). A test on an alloy, the alloy anchor (E221), decides whether the rung can be trusted. Until then it records without gating (passing or failing anything), and the search walks do not climb it.

Blind toFast paths (short-circuit diffusion) along grain boundaries, where crystals meet, and along dislocations, line defects in the crystal. This rung is bulk self-diffusion, through the crystal’s interior only. Transport helped by radiation or by interstitials, extra atoms squeezed between the sites.

Rung 4 · 1.2–1.4 h · typical 1.4 h per 54-atom cell. Locally a cell runs on six MPI ranks (parallel processes) and needs 11–13 GB of scratch disk; E223’s three cells took 1.2–1.4 h each on an A100 GPU (runs/e223_dft_cells/*/pw.out)

First principles

Does a full quantum-mechanical calculation confirm the energies the cheaper rungs use?

Rung 4 is density functional theory (DFT), a quantum-mechanical calculation of the electrons. Every cheaper rung is checked against it. It takes more than an hour per cell, so it runs only on picked compositions.

Precisely: do MACE and DFT put the same gap between this composition and its reference?

The program is Quantum ESPRESSO. The cells are fixed, 54 atoms each, with the atoms on ideal bcc sites. The lattice constant is the Vegard constant: the composition-weighted average of the pure elements’ lattice constants. Nothing relaxes. The settings:

  • PBE, the approximation for how electrons interact, the same as in the training data;
  • 60/720 Ry, the cutoffs that set how finely the wavefunctions and the density are described (Ry, the rydberg, is an energy unit);
  • k 4×4×4, the grid on which the electrons’ waves are sampled;
  • Marzari–Vanderbilt smearing of 0.02 Ry, which smooths how electrons fill their states.

It is the top rung (terminal): nothing sits above it.

dft diagram
Rung 4’s calculation, a loop: the electron density sets a potential, the potential gives the electron states (orbitals), and those give a new density. It repeats until the density stops changing.

Known answerThree independent cells of B2 Mo–Ta agree to 0.01 meV/atom. One has 16 atoms at k 7×7×7; two have 54 atoms at k 4×4×4. The same crystal at the wrong lattice constant was 41 meV/atom off. So the store of DFT results now refuses a lattice-constant mismatch that is not stated (E211). A 2×2×2 k-grid is refused outright: it was 56 meV/atom off at 16 atoms.

Blind toMagnetism, deliberately. The training data were computed without spin (not spin-polarised), and a spin-polarised rung 4 would not be comparable with rung 0. So chromium’s reference is non-magnetic bcc Cr. Configurational entropy. Finite temperature: these are static energies, with the atoms held still.

Rung 1, watched

What does ordering look like as an alloy cools? Most finds stop at rung 1 (19 of E240’s 32), so here it is at work. Each frame is the arrangement of atoms the sampler saved at the end of one temperature step, as it cooled. Below the frame are two measurements from that step. One is the energy. The other is α1, the Warren–Cowley order parameter for nearest neighbours: 0 for a random arrangement, negative when unlike neighbours are preferred. All four runs use v5, the model before embedding 6, because no run on embedding 6 saved its arrangements. E242 (4) repeated the Mo–Ta known answer on embedding 6: 1149 ± 114 K, α1 = −1.00.

The temperature axis. Three runs, on 432 sites, come from the ladder’s own Metropolis sampler, which samples at the temperature it reports. The fourth run comes from pyeCE’s sampler, which runs at twice its stated (nominal) temperature. Its acceptance test adds up per-atom cell energies, so it sees only half of each energy change (E215). The factor was measured as 2.000 ± 0.03 (E215c). That run’s axis is converted to the physical temperature, with the nominal value shown beside it. pyeCE’s short-range-order column is not used here: it holds (α + σ)/2, where σ is the spread of α, rather than α itself. Instead, α1 for that run is computed from its saved arrangements. The dashed line marks the steepest fall of E on the recorded temperature grid: rung 1’s dE/dT estimate, on cooling. The band is the 90–1000 K service window. With reduced motion switched on, the page shows the lowest temperature.

Rung 0: why embedding 6 ships

Which energy model should rung 0 use? It must get two kinds of arrangement right, random and ordered, and the first two candidates each got one. v5 describes each element with three numbers (embedding dimension 3). It gets random solutions right but under-binds ordered ones, putting them too high in energy: B2 Mo–Ta sits 29.5 meV/atom above DFT. Embedding 9 fixes the ordered states. But it over-binds random solutions of compositions it has not seen, putting them too low. It misses by about 13 meV/atom on two of E223’s three random cells. Its seed spread, the disagreement between its three fits, is only 1.2–1.8 meV/atom, so nothing flags the miss (E229, E223). E237 tried the embedding sizes in between; embedding 6 is the one that holds both. The plot moves between the three models on the same fourteen DFT cells. Each point is one cell, with DFT’s energy along the bottom and the model’s up the side. A perfect model would put every point on the diagonal.

ordered cellrandom cellrandom cell of an E223 findbars: spread over the three fits (seeds)

Each model value is the mean of its three fits. It is DFT plus the model’s miss on that cell (the per-cell residual). The residuals are recorded in runs/e229_report.txt (v5, embedding 9) and runs/e237_report.txt (embedding 6), to 1 meV/atom. E223’s cells come from runs/e238_report.txt. DFT values come from data/dft/*.extxyz at 60/720 Ry. Cell list: scripts/validate/score_unbiased_qe.py. In the table, MAE is the mean absolute error, the average size of the miss; the limits (≤) are bars set in advance.

The model must also carry chemistry it has not seen. The test is 48 held-out RHEA Mo–Nb–Ta–W rows: cells kept out of the fit and used only to score it. There embedding 9 loses, at 11.0 meV/atom against v5’s 7.1 (the held-out column above). E238 then refitted embedding 6 on corrected labels. Every version (arm) fitted on the new labels over-bound the search’s random cells. They missed by 13.6–19.6 meV/atom, against a bar of 8. A diagnostic check traced the problem to the median Ti, Cr and Hf reference energies. So the pre-registered fallback holds: rungs 0 and 1 run on E237’s embedding-6 ensemble on the original labels. E239 is measuring those references with DFT anchor cells, cells from the training data recomputed in DFT.

The ladder at work: 32 finds

Where do the search’s finds stop? E223 ran one search with two strategies (arms), each run twice from different random starts (two seeds each). One arm is the fly-head: a swarm of searchers steered by the wiring diagram of a fly’s brain. The other is an elitist baseline, which keeps the best finds and mutates them. The search sorts its finds into niches, bins of valence electron concentration and atomic size mismatch, and keeps the best of each. E223 walked its 32 distinct niche leaders up rungs 0, 1 and 2. Rung 3 is not walked while its trust flag, the switch that lets it pass or fail, is off. Rung 4 is launched separately, on picks. The same 32 finds were walked twice. Hover or tap a dot for its record.

Under E223 the ladder ran on v5, and each cooling run ended at 100 K. 24 finds had no verdict, and none passed. Here p is the ladder’s probability that a composition stays one phase across the window; the search counted a pass at 0.8. The highest was Mo42Ta29Ti29, at p 0.57, although its peak at 896 ± 227 K lies inside the window. The ladder allows for its temperatures running up to 35 % low (scale in forager/ladder/rungs.py), and with that allowance a little over half the chance lands above 1000 K. Under E240 every find has a verdict:

  • 19 order inside the window;
  • 2 fail at rung 2: Mo79Hf19 and Mo70Hf20Ti10, with MACE driving forces of +61 and +35 meV/atom above the hull;
  • 11 pass, with p 0.50–0.77: neither channel shows a transition down to 60 K.

Five of the 19 fail on the fluctuation channel alone, because their slope channel peaked only at its floor. Failing them is the safe direction, but they are not established as ordering either (E242 (4)).

Which of the 11 should go to DFT first? E241 put three independent votes to them:

  • MACE’s rung-2 driving force;
  • the empirical rules on atomic size mismatch (δ) and on valence electron concentration (VEC, the average number of outer electrons per atom);
  • v5’s own verdict from E223.

Ten earned all three. The survivor is Mo51Ti38W3Ta3:

pMACE driving forcesize mismatch δVEC
0.771−110 meV/atom3.5 %5.13

E242 computes its ordered and random cells in DFT. It was pre-registered on 2026-09-23 and has no result yet.

All 32 finds, as recorded
rankcompositionsearch armrung-0 pE223 rung 1E223 pE240 rung 1E240 pMACE driving force, meV/atomE240 stops at
0Mo64Ti23Nb6W6elitist0.99none above 100 K—208 ± 114 K0.12−150rung 1
1Mo58Ti26W8Nb8elitist0.99none above 100 K—331 K (fluct.)0.00−103rung 1
2Mo58Ti21Nb10W5elitist0.98none above 100 K—339 ± 227 K0.14−180rung 1
3Mo56Ti14Nb13W11Ta5elitist0.98454 ± 227 K0.09495 ± 114 K0.00−96rung 1
4Mo53Ti33W14fly0.98none above 100 K—none to 60 K0.76−93passes
5Mo52Ti20Nb14W13fly0.98none above 100 K—402 ± 114 K0.00−87rung 1
6Mo59Ti30W11fly0.98547 ± 227 K0.12none to 60 K0.77−115passes
7Mo51Ti28W16Nb5fly0.98none above 100 K—397 ± 114 K0.00−112rung 1
8Mo62Nb23W13fly0.98none above 100 K—none to 60 K0.73−75passes
9Mo65Nb20W13elitist0.98none above 100 K—353 K (fluct.)0.00−96rung 1
10Mo66Nb22W11fly0.98none above 100 K—none to 60 K0.71−68passes
11Mo58Ti42fly0.97none above 100 K—none to 60 K0.74−83passes
12Mo56Ti41fly0.97none above 100 K—none to 60 K0.75−84passes
13Mo48W29Ta12Nb8fly0.96424 ± 227 K0.10211 ± 227 K0.27−93rung 1
14Mo56Ti19W14Hf5elitist0.96603 ± 227 K0.15582 ± 114 K0.02−82rung 1
15Mo51Ti38W3Ta3 ★elitist0.96none above 100 K—none to 60 K0.77−110passes
16Mo52W25Ta22elitist0.96none above 100 K—364 K (fluct.)0.00−124rung 1
17Mo67Nb12Ta10Cr6W5elitist0.95none above 100 K—none to 60 K0.71−68passes
18Mo52Ti38Nb6elitist0.95none above 100 K—198 ± 20 K0.00−104rung 1
19Mo54Ta16W13Ti7Hf5elitist0.95569 ± 227 K0.12626 ± 114 K0.04−60rung 1
20Mo42Ta29Ti29fly0.95896 ± 227 K0.57795 ± 227 K0.41−114rung 1
21Mo72Nb15Ta8Cr5fly0.94none above 100 K—380 K (fluct.)0.00−47rung 1
22Mo58Ti34Hf8fly0.94none above 100 K—none to 60 K0.50−18passes
23Mo83Hf9Ti7fly0.93none above 100 K—575 K (fluct.)0.00−11rung 1
24Mo65W16Cr7Ti6Nb4elitist0.93none above 100 K—none to 60 K0.73−77passes
25Mo64Ti10Cr10Ta10fly0.91none above 100 K—700 ± 227 K0.22−48rung 1
26Mo69W20V11fly0.90none above 100 K—181 ± 20 K0.00−50rung 1
27Mo69W25Nb6fly0.89none above 100 K—none to 60 K0.64−46passes
28Mo41Ti25Ta10Nb8W7Hf4Zr4elitist0.87544 ± 341 K0.25602 ± 341 K0.27−65rung 1
29Mo48Ti25Nb18Hf9elitist0.87none above 100 K—206 ± 341 K0.27−25rung 1
30Mo79Hf19fly0.80none above 100 K—none to 60 K0.09+61rung 2
31Mo70Hf20Ti10fly0.801291 ± 227 K0.24none to 60 K0.19+35rung 2

From runs/e240_rescored/summary.jsonl, runs/e223_walk/ and runs/e241_agreement.json. p is the ladder’s probability that the composition stays one phase across 90–1000 K; rung-0 p is the screen’s alone. The rung-1 columns give the transition temperature. “none above 100 K”: E223’s runs stopped at 100 K. “fluct.”: the fluctuation channel’s peak, used because the slope channel peaked only at its floor. ★ E241’s survivor.

What no rung sees

Magnetism is the largest hole, and every rung shares it. The energy model, MACE and the DFT frame all ignore electron spin: they are non-spin-polarised. Yet three magnetic or structural changes fall inside the window (E118):

  • chromium’s Néel point (311 K), below which it is antiferromagnetic;
  • nickel’s Curie point (627 K), below which it is ferromagnetic;
  • cobalt’s hcp–fcc change (695 K), between two close-packed crystal structures.

The spin correction is the energy with spin minus the energy without, each measured at the element’s own minimum (E216):

spin correctionCrNiCoFe
meV/atom−12.4−61.1−172.6−460.2

So the non-spin frame is a 12 meV/atom correction on chromium, but a different metal for the three ferromagnets, Ni, Co and Fe. A composition carrying Ni, Co or Fe would need spin-polarised references from its first cell.

Second, the off-lattice hull holds no compound of three or more elements. Suppose a ternary (three-element) compound is more stable than the best binary mixture. Then the hull sits too high, and a find looks more stable than it is. That is the same one-sided, optimistic error as the Laves phases that put seven early compositions 98 to 185 meV/atom wrong (E54, E56). Third, rung 1 sees only order that a 432-site cubic cell can hold, on ideal sites.

Data on this page: runs/e216_v5_metropolis, runs/e210b_scorecard, runs/e211_ground_states/MoTa/start0 (lattice); runs/e240_rescored/summary.jsonl, runs/e240_walk, runs/e223_walk, runs/e241_agreement.json (walk); data/dft, runs/e229_report.txt, runs/e237_report.txt, runs/e238_report.txt (parity), exported by scripts/site/export_ladder_data.py. Rung costs, errors and blind spots: forager/ladder/spec.py. Narrative: EXPERIMENTS.md.

Forager · darth_hidious · MIT licence for the code. Connectome: MaleCNS v1.0 (FlyEM / HHMI Janelia, University of Cambridge, MRC LMB, Google Research), CC BY 4.0. Energy model: an embedded cluster expansion (pyeCE) trained on the RHEA DFT database. Off-lattice screen: MACE-MP-0. DFT: Quantum ESPRESSO. Formal checks: Lean 4 and Mathlib. Structures rendered with OVITO; line drawings made with MetaPost and the fiziko library.
Built with PRISMWebsite and visualizations made using Claude