Forager · what this project added

What we added

Forager stands on other people's physics: the cluster expansion, pyeCE's embedded version of it, the Metropolis sampler, the MACE potentials, Quantum ESPRESSO, the RHEA DFT database and the fly connectome. This page is the other side of that ledger: the ideas this project came up with itself, what each was for, how it was checked, what it bought in a measured number, and whose work it stands on.

The order is by how much of the result depends on each idea, the fidelity ladder first. Where an idea is ours only in how it was applied, the section says so. Each drawing shows the mechanism; the Book has the physics behind it, and every E-number links to its entry in the log.

Where the ideas came from. Six of the questions that shaped this page were the operator's, marked "(operator)" in the log; each section they shaped says so.

The rest came out of the work itself, as the log records it.

  1. 01The fidelity ladder, with every bar written first
  2. 02Smearing which atom sits where
  3. 03The sampler that ran at twice the temperature
  4. 04Labels a lattice model can learn from
  5. 05The exact evaluator
  6. 06Two readings of one heat capacity, and a floor below the window
  7. 07Seed spread, measured where it shows
  8. 08The generator, and the control that can sink it
  9. 09Kenyon-cell sparsity, set on the code that fires
  10. 10The survivor, chosen by agreement
  11. 11Proofs tied to the tests of the code

01 · ours

Target set by the operator (E39)

The fidelity ladder, with every bar written first

idea ladder diagram
Five rungs, each asking a question the rung below cannot ask. A candidate climbs only by passing (white); most stop (toned). The card at each rung is its bar, written in the log before the rung runs. Rung questions from forager/ladder/spec.py.

The problemA cheap model can score millions of compositions, but it can only be wrong in ways it cannot see. The first survey of the design space passed seven compositions. Off the lattice, every one of them loses to crystals a bcc model cannot represent (E54, E56).

The ideaNo single model gives the verdict. Each candidate climbs a ladder of rungs, and each rung asks a different physical question that the one below cannot: its energy on the bcc lattice (rung 0), whether it orders inside 90–1000 K (rung 1), whether a different crystal off the lattice is lower (rung 2), whether order could form in time (rung 3), and a full DFT calculation on the find itself (rung 4). A candidate goes up only by passing. From E35 on, the log's decisive tests are pre-registered: the prediction, the bar that would falsify it and the decision it drives are written before the run, and the result is scored against them, including when it fails.

What it boughtThe off-lattice rung withdrew all seven of the survey's qualifiers. At 90 K each sat 98–185 meV/atom above the cheapest mixture of competing phases (E56). At HfV₂ and ZrV₂ the C15 Laves phase lies 119–123 meV/atom below bcc, 21 times the expansion's own error (E54). On the current walk, 32 finds entered rung 1 and 11 passed rungs 0–2 (E240). One goes to DFT (E242, running). The bar decides what ships: E238's new labels failed theirs, so the model that ships is E237's.

From the operator
The target the ladder is built around is the operator's. In E39 they set the deliverable as short-range order and order–disorder transition temperatures, not the lowest free energy; that is the question rung 1 asks of every candidate.
Verified by
Each rung must pass a known answer before its verdicts count. Rung 1 must put equiatomic Mo–Ta's ordering inside 500–2600 K; it gives 1149 ± 114 K, against a published cluster-expansion Monte Carlo value of 2020 ± 545 K (E242). Rung 2's hull reproduces all seven Laves phases the literature reports as stable in this system (E56).
Builds on
Tiered screening, with cheap filters ahead of expensive calculations, is the standard high-throughput pattern (Curtarolo et al., Nat. Mater. 12, 191, 2013). Pre-registration comes from the replication literature (Nosek et al., PNAS 115, 2600, 2018). Every rung's model is someone else's: pyeCE, MACE, Quantum ESPRESSO. Ours are the questions, their order, the known answer each rung must pass and the rule that the bar comes first.

02 · ours

Idea: the operator (E34)

Smearing which atom sits where

idea smearing diagram
A site's identity, sharp and smeared, and the energy along a smeared path from one arrangement to its exchange. The curve has the span and bow E36 derived for Ti–W; its direction and the side of the bow are drawn for illustration.

The problemWhich atom sits on which site is a discrete choice. A search can only jump from one arrangement to the next; nothing smooth connects them, so a method that follows a slope has none to follow. The question came up on 12 September, just after a Gaussian process had beaten the fly on the search (E33).

The ideaThe operator's, by analogy with DFT: DFT replaces the sharp step between filled and empty electronic states with a smooth function so that the energy can be differentiated. Smear atom identity the same way. Each site becomes a mixture of two elements, and a parameter t carries the cell from one arrangement to its exchange, so the energy becomes a smooth function between arrangements. On this axis the project's own models line up: t = 0 is one fixed arrangement, and every site smeared to the alloy's overall composition, independently of its neighbours, gives in expectation the random-solution energy that rung 0 reports.

What it boughtThe relaxation works across groups 4 to 6. Seven of eight element-exchange paths are monotone, and every group-4-to-6 pair is: Ti–W spans 16.7 meV/atom, Nb–Mo 28.6. The farther apart two elements sit in the periodic table, the larger the span (correlation +0.69; E36). The prediction written in the log before E34 ran was the opposite, that interpolating Hf into W would be too unphysical to smooth anything, and the log keeps it beside the measurement. The result also withdrew the redirection E33 had proposed, which assumed this was a regime where kernel methods are structurally weak (E34).

Verified by
E34 sampled paths between an arrangement and its exchange on 16 sites, swapping sites in pairs so the composition never changes. A first attempt let each site draw its element independently, did not preserve composition, and measured composition fluctuation instead (thousands of meV/atom); it is recorded as wrong. E36 then derived the curve instead of sampling it. Under independent site occupations the expected energy factorises, and because every cluster touches at most three sites, ⟨E⟩(t) is a cubic in t whatever the cell size. This was checked symbolically, with no sampling error, and it showed E34's apparent turning points were sampling noise. These are two-element paths in a 16-site cell, not the whole nine-element surface.
Builds on
The operator reached the idea independently, and the method is published: Kaappa, Larsen & Jacobsen, Phys. Rev. Lett. 127, 166001 (2021), interpolation between chemical elements, demonstrated on Au–Cu bulk and on Cu–Ni surfaces and clusters, pairs that are neighbours in the periodic table. E34 and E36 extended the test to chemically dissimilar refractory pairs across groups 4 to 6. Smearing occupations is standard practice in DFT.

03 · ours

Scorecard started by the operator (E42)

The sampler that ran at twice the temperature

idea sampler diagram
One swap on the model v5, priced two ways: the true change in total energy and what pyeCE's sampler used in its acceptance test. The bars are to scale (scripts/ordering/pyece_delta_check.py).

The problemFor two days rung 1 put every ordering temperature at about half the published value. The model was blamed, and refits were built to cure it (v6a and v6b, E190b).

The ideaStop asking the model and test the sampler on a problem whose answer is known exactly. Fit the same network to a nearest-neighbour Ising model on bcc (J = 19.12 meV, exact transition 1410 K), give it to pyeCE's sampler, and compare with an independent Metropolis sampler written here. Then read pyeCE's source for a convention that would produce a factor. It was there: the model gives an energy per atom for each two-site cell, the sampler adds cells to price a swap, so it sees half the true change and samples at twice the temperature it was asked for. The same aggregation turns its short-range-order column into the average of α and its own standard deviation, (α + σ)/2. That can be decoded exactly for an ordering pair (α < 0) and not at all for α > 0, so the decoder refuses rather than guesses.

What it boughtOrdering temperatures became comparable with the literature. On the physical axis, v5 put Mo–Ta at 910–1037 K by every estimator (published 2020 ± 545 K) and MoNbTaVW at 697–702 K (published 742 and 750 K). Against Sobieraj et al.'s table, four of five systems land within 22 %; the fifth is a Cr–Ta disagreement that DFT then settled (E222). The refits' starting premise, "rung 1 reads half the published value", was the sampler's factor, not the model's.

From the operator
The external scorecard this section scores against began with the operator's request in E42 to reproduce a result published by a professor on the project, Sobieraj et al. 2020, on Ta–Ti–V–W. It came out at 383 ± 75 K against their 500 K, with the same strongest ordering pair, Ta–W.
Verified by
The exact Ising model came out at 612 K nominal, against 1300–1340 K from the independent sampler (E215). On a fine grid the factor is 2.000 ± 0.03 (E215c). On pyeCE's own sampler object, v5's true swap change of 538.4 meV was seen as 269.2 meV, a ratio of exactly 0.500. Decoded α agrees with the independent sampler to within 0.007 at four temperatures. Tests pin both conventions (tests/test_pyece_sampler_factor.py, tests/test_pyece_sro_encoding.py), so a future pyeCE that changes them turns a test red. Adversarial code reviews caught an earlier, guessed reading of the SRO column and then confirmed the decoding algebra.
Builds on
The Metropolis algorithm (Metropolis et al., J. Chem. Phys. 21, 1087, 1953); Warren–Cowley short-range order (Cowley, Phys. Rev. 77, 669, 1950); pyeCE and the embedded cluster expansion (Müller & Natarajan, npj Comput. Mater. 11, 60, 2025); and the published numbers it was scored against (Kim & Widom, Phys. Rev. Materials 7, 063803, 2023; Sobieraj et al., PCCP 22, 23929, 2020). The upstream fix is pyeCE's to make. Here every temperature is read as nominal × 2, with the factor stated.

04 · ours

Question from the operator (E26)

Labels a lattice model can learn from

idea labels diagram
Left, a RHEA frame as DFT computed it. Right, what a cluster expansion can represent. The label keeps DFT's energy and removes the shake and the strain with one MACE difference on the same atoms (scripts/expansion/rhea_labels.py). Species and displacements are drawn for illustration.

The problemThe RHEA database was built to train interatomic potentials. Its atoms sit 0.13 Å off their sites on average, worth about 200 meV/atom, and its cells range from 2.665 to 3.818 Å, worth up to about 1700 meV/atom. A cluster expansion sees only which atom sits on which ideal site, and the ordering signal it has to resolve is 10–30 meV/atom. Trained on the raw energies, it fits noise (E150).

The ideaKeep DFT's energy and change only the geometry it describes. For each frame, subtract one difference from a machine-learned potential: MACE on the frame as RHEA has it, minus MACE on the same atoms at ideal bcc sites and their relaxed lattice constant. MACE enters only as that difference on one structure, never as an absolute energy. The pure-element references are fitted in the same convention; zirconium has no usable bcc cell in RHEA at all. Each row carries its own checks: whether the volume minimum was bracketed, and, where it applies, a MACE-free estimate from RHEA's own forces.

What it boughtFormation energies centred at 0.0 meV/atom (sd 99) on the first 130 frames, against +622 before the volume term was added (E150). The model that ships at rungs 0 and 1 is trained on these labels. It misses the eleven independent Quantum ESPRESSO cells by 7.5 meV/atom on average and the search's own three random cells by 6.1 (E237). Refinements are built and tested but not shipped yet: the true lattice minimum per frame (the original volume fit sat 17 meV/atom high, E228), non-cubic ordered cells (2,031 of 2,059 now map, E234) and a gate on far-displaced sites. With median pure references they over-bind random cells (19.6 against 6.1, E238). E239 is now measuring those references directly.

From the operator
In E26 the operator asked why the energy model was fitted to a potential and not to RHEA's DFT in exactly this chemistry. The answer recorded then was that RHEA's frames are the wrong shape for a lattice model: displaced atoms cost 125.5 meV/atom with 36.5 of scatter. The label correction is what gave them the right shape, and rung 0 is now trained on RHEA's DFT.
Verified by
MACE is not soft on these frames: force slope 1.014 and virial slope 1.072 against DFT, an implied bias of −0.2 meV/atom over 150 frames (E228). Four of RHEA's deepest ordered cells, recomputed in Quantum ESPRESSO, land within 13 meV/atom of their labels (E232); B2 Mo–Ta's label is −183.9 meV/atom against QE's −184.5. The label algebra is stated in Lean and tied to tests (the last section).
Builds on
The RHEA DFT database (Byggmästar, Lopes, Fan & Ala-Nissila, arXiv:2603.04147, 2026; Zenodo 10.5281/zenodo.18863415); the MACE foundation potentials (Batatia et al., arXiv:2401.00096); and the cluster expansion's definition on the ideal lattice (Sanchez, Ducastelle & Gratias, Physica A 128, 334, 1984). The per-element anchor correction E239 applies is Materials Project-style practice, named as prior art in E164, and is not ours.

05 · ours

The exact evaluator

idea exact diagram
The same trained network, evaluated two ways. Top, pyeCE's forward pass: each cluster a swap touches is gathered and priced one at a time. Bottom, ExactECE (forager/ladder/ece_distill.py): correlation tensors per orbit, with the embedding folded into the first layer. Swaps per second at embedding 9, one thread, 432 sites (E237).

The problemRung 1 prices about 1.8 million swaps per composition (E238). pyeCE's forward pass managed 28 swaps a second at embedding 9, the size that got Mo–Ta right, so one rung-1 walk would take about 15 hours (E233). Accuracy and cost were pulling in opposite directions.

The ideaRe-evaluate the model instead of approximating it. The eCE energy is a small network applied to correlation functions, and those are sums over clusters. Written as one correlation tensor per orbit, with the linear embedding folded into the first layer, the same numbers come out of a few tensor contractions instead of a loop that gathers each cluster.

What it bought1,519 swaps a second against 28 at embedding 9 (54×), and 2,788 against 1,009 for v5 (E237). Rung-1 cost stopped depending on the embedding size, so the shipped model (embedding 6) was chosen on accuracy alone. Rung 1 now samples the mean of the three-seed ensemble.

Verified by
Parity with pyeCE's own forward pass: over 2,304 swaps, a mean difference of 0.0012 meV and a maximum of 0.0071 (E237). An adversarial review on seven models found per-cell differences of at most 6.3 × 10⁻⁷ eV and no measurable drift over more than 1,300 accepted swaps. Rung 1 on the same model gives 1148 ± 114 K against 1175 ± 114 K through pyeCE. The fold is proved in Lean and tested. A tiling known answer covers pairs longer than half the cell (1.5 × 10⁻⁵ meV/atom). A fitted surrogate was tried and failed parity (55.7 meV RMS per swap), although its transition landed at 1170 K. That is why parity, not the transition temperature, is the acceptance test.
Builds on
The embedded cluster expansion and its network are Müller & Natarajan's (npj Comput. Mater. 11, 60, 2025; the pyeCE software), and the correlation-function form of the energy is Sanchez, Ducastelle & Gratias's (1984). Ours are the re-evaluation and its checks.

06 · ours

Two readings of one heat capacity, and a floor below the window

idea channels diagram
The heat capacity read two ways while the model cools. Solid: the size of the energy's fluctuations. Dashed: the slope of the energy. The shapes are schematic; the positions are those E240 recorded for Mo₅₂W₂₅Ta₂₂: the fluctuation peak at 364 K and the slope maximum at the 60 K floor.

The problemA transition is read from where the heat capacity peaks as the model cools. Rung 1 used to stop at 100 K while the window starts at 90 K, so an alloy that never orders could not pass by construction. 24 of E223's 32 finds, including its top three, were left without a verdict. A single reading can also be fooled when the sampler stops moving at the cold end.

The ideaRead the heat capacity two ways at once: from the slope of the energy with temperature and from the size of the energy's fluctuations. In equilibrium the two are the same number. Where they disagree, the sampler is out of equilibrium or not at the temperature it reports. Keep cooling to 60 K, below the window, and score "no transition down to the floor" as a bounded pass instead of a gap.

What it boughtAll 24 of E223's no-verdict finds now carry one, and 11 of 32 pass rungs 0–2 (E240). Checking the chosen survivor's curves before any DFT time was spent showed that it, Mo₅₂W₂₅Ta₂₂, was one of five finds whose slope maximum sat at the floor while the fluctuations peaked inside the window, at 331–575 K. The pick was voided and the scoring fixed.

Verified by
The ratio of the two readings was 2.00 in every pyeCE sweep and 0.94 for the exact control, the signature of the 2T sampler, and 2.01 again on E215c's fine grid. The known answer E242 exposed the opposite artefact: a spurious fluctuation peak at 80 K while the true transition sits at 1149 K. So the fluctuation reading is used only where the slope reading finds no peak, and the five finds above are reported as "not established", not as ordering. The fixes carry tests; all 32 finds were re-scored without re-sampling.
Builds on
That the heat capacity equals both the energy's temperature slope and its fluctuations divided by kT² is textbook statistical mechanics (Landau & Binder, A Guide to Monte Carlo Simulations in Statistical Physics, Cambridge). Sobieraj et al. define their published transition as an inflection of the mixing enthalpy, which is the same heat-capacity peak, and that is what made their table comparable (PCCP 22, 23929, 2020). Ours are the use of the pair as a running check and the floor rule.

07 · ours

Question from the operator (E40)

Seed spread, measured where it shows

idea seeds diagram
One model (v5's labels and recipe) fitted with three random seeds and scored two ways. Each sphere is a seed. On 48 held-out rows of the training database the three agree; on eleven DFT cells no model was steered toward, they do not.

The problemEvery refit was judged on 48 held-out RHEA rows. Refit the same model with two more random seeds and that number barely moves: 7.2, 7.4 and 7.5 meV/atom.

The ideaTrain every candidate model as at least three seeds and judge the ensemble on DFT that no model was chosen or steered by: eleven Quantum ESPRESSO cells, and later the search's own random cells. The spread across seeds is reported as the model's error bar on each cell.

What it boughtThe same three fits miss the eleven cells by 22.9, 16.6 and 23.8 meV/atom. The seed spread averages 8.5 meV/atom per cell and reaches 25 on the most strongly ordered one (log, 22 Sept., before E229). Judged this way, embedding 5 was not shipped although it passed its own bars: it missed Mo₇₀Hf₂₀Ti₁₀ by +23.9 meV/atom with a seed spread of only 2.9. Embedding 6 held both tests (E237).

From the operator
In E40 the operator required that the circuit learn uncertainty properly. There, a bootstrap ensemble's spread proved worse than a constant (interval width 1.763 against 1.124) because the error was a reproducible bias, not a variance; conformal prediction made every signal honest (after Hu, Musielewicz, Ulissi & Medford, Mach. Learn.: Sci. Technol. 3, 045028, 2022). That is the standard this section holds an error bar to: it counts only once it is checked against errors the model did not see.
Verified by
The rule caught two confidently wrong models. Embedding 9 over-binds two of the search's three random cells by about 13 meV/atom with a seed spread of 1.2–1.8 (E223, rung 4), and embedding 5 is the case above (E237). In both, the spread alone would not have flagged the error; the unsteered DFT cells did.
Builds on
Ensembles of independently trained networks as an uncertainty estimate (Lakshminarayanan, Pritzel & Blundell, NeurIPS 2017). Müller & Natarajan already report that eCE validation errors vary across random initialisations (npj Comput. Mater. 11, 60, 2025). Ours are finding that the held-out metric was blind to the spread, and the rule that every refit is judged as an ensemble on DFT no model chose.

08 · ours

Question from the operator (E32)

The generator, and the control that can sink it

idea generator diagram
Walkers step over compositions (a three-element slice of the nine is drawn). A readout of the fly's brain scores each proposed step and learns from the ladder's rewards. E224's control drives the same swarm with the wiring rewired so that every neuron keeps its number of links. Positions are illustrative.

The problemThe ladder can only judge what it is given. Something has to propose alloys across a nine-element space, and the claim that a fly connectome helps has to survive controls that could sink it.

The ideaA swarm of walkers moves over compositions. Each proposed step is scored by a readout of the male fly's whole central nervous system (the MaleCNS connectome) driven by the composition, and the readout learns from the rewards the ladder pays. The claim is tested against arms that share everything but one ingredient: the real wiring, a degree-preserving rewire, a random head of matched sparsity, a frozen twin, keep-the-best-and-mutate (elitist), elitist with the swarm's diversity devices, a ridge readout driving the same swarm, and MAP-Elites. Arms are paired by seed, five seeds each, with the bar written first (E224).

What it boughtAt three seeds on the pure reward, the swarm found 299 ± 31 distinct qualifying alloys against elitist's 238 ± 14 (+61, two seed spreads), at the same rate (AUC_Q +2.3, a tie; E195c, E201b). Its readout does learn the reward (E201). An adversarial review traced the advantage to width, not depth: a proposal filter, restarts and 24 walkers against 8. That is why E224 carries an elitist arm with the same devices.

E224 · in progress, read at build time

Distinct qualifying finds in 800 verifications, real wiring / rewire, by seed: 333 / 301, 278 / 332, 333 / 315, 15 / 240, 203 / 334 (A − B: +32, −54, +18, −225, −131). 2 of 5 differences are positive and the median is −54. On these five pairs the bar is not met: the real wiring does not beat the rewire. Seed 3 was a cold-start stall for the real and the random heads. The remaining arms (frozen twin and the four baselines) have 4 of 25 seeds recorded; no verdict card yet.

runs/odt_map/stage_b_e224_<arm>_s<seed>.json

From the operator
In E32 the operator asked why the number of proposals per round was the experimenter's choice. The answer was a stopping rule from foraging theory, Charnov's marginal value theorem, written into the swarm's walker (forager/search/forage.py): the same mean result, a tighter spread, one fewer hand-set number. Today's tests run at a fixed budget of 800 verifications, so the rule is not what E224 measures.
Verified by
E224 itself: the real wiring must beat the rewire on all five seeds with a median of at least 31 finds. The review's prior that it would was below 15 %.
Builds on
The MaleCNS v1.0 connectome (Berg et al., bioRxiv 10.1101/2025.10.09.680999; FlyEM at HHMI Janelia, Cambridge, MRC LMB and Google Research); degree-preserving rewiring as the null (Maslov & Sneppen, Science 296, 910, 2002); MAP-Elites (Mouret & Clune, arXiv:1504.04909, 2015). Ours are the use of a connectome as the scoring head of a search over alloys, and the control set.

09 · ours

Kenyon-cell sparsity, set on the code that fires

idea kc diagram
Twenty Kenyon cells, their raw input (bars) and the threshold (dashed). White cells fire after inhibition; toned cells are silent. Top: the threshold set on the raw input. Bottom: set on the cells that actually fire (forager/brain/mushroom.py, _match_evoked_sparsity). Bar heights are illustrative.

The problemIn the fly's mushroom body only a few per cent of Kenyon cells fire for an odour, and learning happens only at synapses from cells that fired. The model set its threshold so that 10 % of the raw input cleared it. Recurrent inhibition then silenced most of what cleared, and on the near-pure compositions a corner-seeking search visits, 0.07 % fired: three cells of 4,064 (E61).

The ideaSet the threshold by bisection against the quantity actually specified, the fraction of cells that fire after inhibition, measured on probe compositions.

What it boughtThe learning fly then beat its own frozen twin on every seed: +4.68 ± 1.40, Wilcoxon p = 0.0039, 8 of 8 seeds, against +0.45 ± 0.51 (p = 0.40) with the starved code. The result repeated on eight fresh seeds, +5.02 ± 0.90 (E61). It was measured before the 19 September fix that stopped the reward depending on process history, so it is a within-run comparison, not an identical-reward one.

Verified by
The probes fire at 10.04 % after calibration. The difference between compositions in the Kenyon layer rose from 1.43 to 4.99, and the held-out correlation between the circuit's score and its reward from 0.291 to 0.560 (E61).
Builds on
Sparse Kenyon-cell coding held in place by feedback inhibition from the APL neuron (Lin et al., Nat. Neurosci. 17, 559, 2014). Neither the gain nor the threshold is in the connectome; setting them against the evoked code is ours.

10 · ours

The survivor, chosen by agreement

idea agreement diagram
The eleven finds that passed rungs 0–2 (E240), placed by three independent votes (runs/e241_agreement.json). Ten carry all three. The highest-scoring find carries two. Placement within a region is arbitrary.

The problemAfter E240, eleven finds passed rungs 0–2 and one DFT run was affordable. The top of one model's ranking is where that model is most likely to be extrapolating.

The ideaPick where independent models agree, not where one model is most confident. There are three votes, none from the shipped model: MACE's off-lattice hull (no cheaper mixture of crystals), the empirical rules from tabulated element data (size mismatch δ < 6.6 %, valence electron count below 6.87), and v5, a different model on different labels, finding no transition inside the window. The rule is most votes, then the highest pass probability.

What it boughtTen of eleven eligible finds carry all three votes. The single highest-scoring find, Mo₅₉Ti₃₀W₁₁ (p 0.774), lost v5's vote because v5 finds it ordering inside the window. The pick is therefore Mo₅₁Ti₃₈W₃Ta₃ (p 0.771, 110 meV/atom below its cheapest competing mixture, δ 3.5 %, VEC 5.13; E241). Its DFT test is running: an ordering energy of at most 5 meV/atom confirms it, above 15 fails it (E242).

Verified by
The rule was written before E240 finished (E241). What it chose is being tested by the rung it was chosen for.
Builds on
The principle is borrowed. The log names Blalock … Romero (bioRxiv 2026.09.15.751856), a preprint that could not be retrieved to check here. The same lab's peer-reviewed work designs against the prediction that most of an ensemble agrees on (Freschlin, Fahlberg, Heinzelman & Romero, Nat. Commun. 15, 6405, 2024). The rules are Yang & Zhang's δ ≤ 6.6 % (Mater. Chem. Phys., 2012, doi:10.1016/j.matchemphys.2011.11.021) and Guo et al.'s VEC < 6.87 for bcc (J. Appl. Phys. 109, 103505, 2011). Ours is only the application to choosing one alloy for DFT.

11 · ours

Proofs tied to the tests of the code

idea proofs diagram
One identity, checked twice. It is proved in Lean 4 over the reals (formal/Temperature.lean) and tested on the production function that computes it (tests/test_formal_temperature.py). Every theorem and its test are listed in formal/README.md.

The problemThe ladder's numbers rest on a handful of identities: the 2T convention, detailed balance of the swap move, the SRO decode, the Warren–Cowley bound, the label algebra and the evaluator's fold. A proof about an idealised formula can hold while the code does something else, and a test can pass on the wrong formula.

The ideaState each identity in Lean 4 with Mathlib, over the reals with explicit hypotheses, and tie each theorem to a pytest of the production function that computes it, so the proof and the code answer to the same statement. An audit test refuses sorry, admit, new axioms and native_decide.

What it boughtSix modules are checked: Temperature, DetailedBalance, SroDecode, WarrenCowley, VolumeMiss and FeatureFold. No proof found the code disagreeing with the mathematics. One proof explains why a check matters: a volume minimum found at the edge of the searched bracket says nothing about what lies beyond it, which is why every label carries a flag for that case. All 5,836 labels of E238 have an interior minimum.

Verified by
An adversarial review of the statements checked their fidelity to the code, checked for vacuity with compiled witnesses, and mutated the code (4 of 4 mutations caught). It found two overclaims, both fixed (E238).
Builds on
Lean 4 (de Moura & Ullrich, CADE 2021) and Mathlib (The mathlib Community, CPP 2020). The theorems are standard mathematics; the tie to the production code is ours.

Considered and left off

Two candidates are not on this page. The spin-aware DFT retry sequence (staged smearing, then a restart at production smearing) is not ours: the retry handlers are aiida-quantumespresso's, ported (log, 22 Sept.), and a staged smearing restart is ordinary practice. Its known-answer run has not been written up in the log yet. The anchor cells that pin the label frame's pure-element references (E239) use a Materials Project-style per-element correction, named as prior art in E164, and the run has no result yet.

Built with PRISMWebsite and visualizations made using Claude