Notebook · 23 September 2026
Forager
Is there a refractory alloy that stays one body-centred cubic phase — no ordering, no decomposition — at every temperature from 90 to 1000 K? Forager searches for one among the nine elements its models are validated on: Cr, Hf, Mo, Nb, Ta, Ti, V, W and Zr.
Two machines work on the question together. A generator — a swarm of walkers moving through the space of compositions, steered by a learned readout of the male fly's whole connectome and taught by the rewards it is paid — proposes alloys. A fidelity ladder of five rungs decides, at rising cost, whether each proposal meets the requirement: a cluster expansion trained on density-functional theory, Monte Carlo for the ordering temperature, a machine-learned potential for competing phases, kinetics, and density-functional theory itself.
Where it stands
The first search to be walked up the corrected ladder sent 32 of its finds up the rungs (E223). Re-walked on the shipped model, after two defects in how the ladder treated alloys that show no transition down to its coldest temperature were found and fixed, 11 of the 32 pass rungs 0–2 (E240). Three independent models — the MACE potential, the empirical size and valence-electron rules, and the previous cluster expansion — agree on one of them, Mo₅₁Ti₃₈W₃Ta₃ (E241). Whether it orders at all is now one comparison at first principles: its model-preferred arrangement against a random one.
E242 · pending
Does Mo₅₁Ti₃₈W₃Ta₃ order? The shipped model's lowest-energy 54-atom arrangement of the survivor and a random arrangement of the same composition are being computed with DFT at the ladder's standard (60/720 Ry, k 4 × 4 × 4). Written before the run: an ordering energy of at most 5 meV/atom confirms the find as the ladder's first full pass; 5–15 leaves it not established; above 15 the find fails. The model's lowest state is only the lowest the model can find, so a small DFT ordering energy cannot rule out a different order that DFT prefers; the entry says so.
Already passed: the known-answer check on the shipped model's ordering rung — equiatomic Mo–Ta orders at 1149 ± 114 K against a published CE + Monte Carlo 2020 ± 545 K.
What we added
The cluster expansion, its embedded form, Metropolis sampling, the MACE potential and the DFT code are other people's work. What this project added is how they were checked, corrected and put together:
- An exact, faster evaluator. pyeCE's energies re-derived in numpy: they agree to within 0.023 meV per swap, and sample 54 times faster at the largest embedding.
- The sampler's factor of two. Found on a transition known exactly, confirmed by a second sampler written from scratch (E215).
- Labels on the ideal lattice. Each DFT frame's rattle and strain removed, so a lattice model can learn from it (E150, E151).
- Uncertainty from seed ensembles. The standard held-out score could not see an 8.5 meV/atom spread between seeds; three-seed ensembles can.
- A survivor by agreement. Three independent models vote, instead of one learned model extrapolating (E241).
The generator

Twenty-four walkers move through the simplex of compositions. Each proposal is encoded as a sensory input — broadcast by a fixed random map onto all 17,479 afferent neurons of the MaleCNS connectome (164,506 neurons, 25.1 million connections) — propagated for four recurrent steps, and scored by a linear readout over the output neurons of the mushroom body, the lateral accessory lobe and the descending neurons. The readout is learned online from the rewards the ladder's cheapest rung pays; nearly all of its weight settles on the mushroom body's outputs (E161, E163).
What it has shown, on a reward that ranks: the readout learns the reward — score–reward correlation +0.90 against −0.03 for a frozen twin (E201) — and the swarm finds a quarter more distinct qualifying alloys than an elitist search at the same rate, 299 ± 31 against 238 ± 14 over three seeds (E195c). What it has not shown: its head-derived gain field was the worst of three tested (E198b), its plastic site in the fan-shaped body never reached the readout (N6), and the fixes meant to make it exploit harder made it worse (E202, E203). The advantage it has is width, not depth. Whether the connectome's wiring contributes anything a matched random graph does not is the question E224 was designed to settle.
E224 · in progress
Does the connectome's wiring matter? Eight arms, five seeds each, paired by seed: the real wiring (A), a degree-preserving rewire (B), a random head of matched sparsity (C), a frozen twin (D), and four baselines — elitist search, elitist with the fly's diversity devices, a ridge readout driving the same swarm, and MAP-Elites (E–H). Seeds recorded so far: A 5, B 5, C 5, D 4, E–H 0 of 20.
| Seed | Real wiring (A) | Rewired (B) | A − B |
|---|---|---|---|
| 0 | 333 | 301 | +32 |
| 1 | 278 | 332 | −54 |
| 2 | 333 | 315 | +18 |
| 3 | 15 | 240 | −225 |
| 4 | 203 | 334 | −131 |
Distinct qualifying finds over 800 verifications. The bar written before the run: all five A − B differences of one sign, with a median of at least 31. 2 of 5 differences are positive and the median is −54. On the five pairs recorded that bar is not met. Seed 3 was a cold-start stall for the real and the random head — qualifying finds per 50-round window 0 / 0 / 1 / 34 and 0 / 0 / 0 / 1 — which a replay is examining. No verdict is drawn until every arm is in.
Read at build time from runs/odt_map/stage_b_e224_<arm>_s<seed>.json; no runs/e224_card.txt yet.
The ladder
Each rung asks a question the rung below cannot, and only survivors climb. The cheap rungs are the same model; the dear ones are independent of it.
| Rung | Asks | Model | Cost | Measured error |
|---|---|---|---|---|
| 0 · screen | How far below its competitors is the random solid solution? | Embedded cluster expansion (pyeCE), three-seed ensemble on RHEA DFT labels | 0.1–0.5 s | 7.5 meV/atom on eleven DFT cells no model chose (E237) |
| 1 · ordering | At what temperature does it order? | Metropolis Monte Carlo on a 432-site bcc cell, 2600 → 60 K, the ensemble mean evaluated exactly | 31 min median per find, rung 2 included (E240 walk logs) | Mo–Ta 1149 ± 114 K; published 2020 ± 545 K (E242) |
| 2 · lattice | Is it bcc at all, once its atoms may move? | MACE-MPA-0 relaxation against a 671-phase off-lattice hull | 8–34 s | 1–4 meV/atom on the gaps the ranking uses (E78) |
| 3 · kinetics | Can atoms travel far enough to reach what the thermodynamics wants? | MACE vacancy formation and 30 migration barriers, gated on the 97.5th-percentile reach | 52 min | Recorded, not trusted, until an alloy anchor lands (E221) |
| 4 · first principles | Does density-functional theory agree? | Quantum ESPRESSO, PBE, 60/720 Ry, 54-atom cells | 3.4–4.5 h | The reference the others are checked against |
What is established
- The ordering rung's original sampler runs at twice the temperature it reports — a factor of 2.000 ± 0.03, found on a transition known exactly and then in the sampler's source (E215, E215c) — and its short-range-order column is (α + σ)/2, not α. Rung 1 now samples on the physical axis, and the experiments that still use the old sampler are read with both corrections.
- On the physical temperature axis, the previous model placed six of six comparable published ordering temperatures within 22 %; the seventh was one pair over-ordered 2.4 times, which DFT confirmed (E210b, E222).
- The shipped model puts equiatomic Mo–Ta at 1149 ± 114 K against a published 2020 ± 545 K (E242), and lies within 7.5 meV/atom of eleven DFT cells no model chose and 6.1 of the search's own random cells (E237).
- The cheapest rung (then v5) predicted the DFT formation energies of three of the generator's top finds to within 8, 3 and 28 meV/atom; the last, Mo₃₆Nb₁₈, it overstated (E194, E196).
- An exact re-evaluation of the model agrees with pyeCE to 0.023 meV per swap and samples 54 times faster at the largest embedding tested.
- Six conventions the ladder's numbers rest on — the temperature factor, detailed balance of the sampler, the short-range-order decoding, the Warren–Cowley bounds, the labeller's volume miss and the exact evaluator's algebra — are proved in Lean 4, each tied to a test of the code that computes it.
What is open
- Does Mo₅₁Ti₃₈W₃Ta₃ order? (E242, pending.)
- Does the connectome's wiring matter? (E224, in progress.)
- The pure-element reference energies of Ti, Cr and Hf, measured now with DFT anchor cells (E239). They move absolute formation energies, not ordering.
- The kinetics rung: MACE gives MoNbTaW a vacancy formation energy of 3.55 eV against 2.48–2.54 eV in the literature, and the alloy anchor (E221) has not landed, so the rung records without deciding.
- Magnetism. Every rung is spin-blind. Chromium, one of the nine, has a Néel transition at 311 K, inside the window (E122); the spin energy the non-spin frame discards is −12.4 meV/atom for Cr, and −61.1 to −467.4 for Ni, Co and Fe (E216).
- Ternary and higher compounds. The off-lattice hull holds 671 phases and none has more than two elements (E147), so every multi-element find sits in its blind spot.
- Five of E240's finds are "not established": their verdict turns on which heat-capacity channel is trusted at low temperature.
How it got here
Forager began as an idea around 1 September; the first files were made on 10 September and the first commit landed on 11 September. It started as a different experiment: a whole-brain network trained by a local forward-forward rule, choosing which of 64 fixed alloys to compute next. It did not choose — the three runs sampled uniformly with a fixed seed — and backpropagation reproduced its picks 39 times in 40 (E1), so the learning rule was not the problem. Forward-forward training was given up. Everything since, including what was withdrawn and why, is in the Logbook; the September runs are kept on the Results page.