Notebook · 23 September 2026

Forager

Is there a refractory alloy that stays one body-centred cubic phase — no ordering, no decomposition — at every temperature from 90 to 1000 K? Forager searches for one among the nine elements its models are validated on: Cr, Hf, Mo, Nb, Ta, Ti, V, W and Zr.

The male fly's central nervous system in three dimensions, drawn from the soma positions recorded in the MaleCNS v1.0 dataset (FlyEM / HHMI Janelia, University of Cambridge, MRC LMB, Google Research; male-cns.janelia.org, CC BY 4.0): 23,995 of its 141,781 somata, sampled by superclass, and in dark stipple 71 of its measured neurons. Drag to turn it, pinch or use + and − to zoom, double-click to set it straight again.

Two machines work on the question together. A generator — a swarm of walkers moving through the space of compositions, steered by a learned readout of the male fly's whole connectome and taught by the rewards it is paid — proposes alloys. A fidelity ladder of five rungs decides, at rising cost, whether each proposal meets the requirement: a cluster expansion trained on density-functional theory, Monte Carlo for the ordering temperature, a machine-learned potential for competing phases, kinetics, and density-functional theory itself.

question diagram
The requirement: one bcc phase at every temperature from 90 to 1000 K — and the two ways to fail it, ordering on the lattice or leaving it.
32finds from the first search walked up the corrected ladder (E223, E240)
11of them pass rungs 0–2, once two floor-rule defects were fixed (E240)
Mo₅₁Ti₃₈W₃Ta₃the survivor: all three independent votes and the highest ladder p (E241)
pendingE242: does it order? Its model-preferred cell against a random one, in DFT
1149 Kequiatomic Mo–Ta on the shipped model, ± 114 K; published 2020 ± 545 K (E242)

Where it stands

The first search to be walked up the corrected ladder sent 32 of its finds up the rungs (E223). Re-walked on the shipped model, after two defects in how the ladder treated alloys that show no transition down to its coldest temperature were found and fixed, 11 of the 32 pass rungs 0–2 (E240). Three independent models — the MACE potential, the empirical size and valence-electron rules, and the previous cluster expansion — agree on one of them, Mo₅₁Ti₃₈W₃Ta₃ (E241). Whether it orders at all is now one comparison at first principles: its model-preferred arrangement against a random one.

E242 · pending

Does Mo₅₁Ti₃₈W₃Ta₃ order? The shipped model's lowest-energy 54-atom arrangement of the survivor and a random arrangement of the same composition are being computed with DFT at the ladder's standard (60/720 Ry, k 4 × 4 × 4). Written before the run: an ordering energy of at most 5 meV/atom confirms the find as the ladder's first full pass; 5–15 leaves it not established; above 15 the find fails. The model's lowest state is only the lowest the model can find, so a small DFT ordering energy cannot rule out a different order that DFT prefers; the entry says so.

Already passed: the known-answer check on the shipped model's ordering rung — equiatomic Mo–Ta orders at 1149 ± 114 K against a published CE + Monte Carlo 2020 ± 545 K.

What we added

The cluster expansion, its embedded form, Metropolis sampling, the MACE potential and the DFT code are other people's work. What this project added is how they were checked, corrected and put together:

  • An exact, faster evaluator. pyeCE's energies re-derived in numpy: they agree to within 0.023 meV per swap, and sample 54 times faster at the largest embedding.
  • The sampler's factor of two. Found on a transition known exactly, confirmed by a second sampler written from scratch (E215).
  • Labels on the ideal lattice. Each DFT frame's rattle and strain removed, so a lattice model can learn from it (E150, E151).
  • Uncertainty from seed ensembles. The standard held-out score could not see an 8.5 meV/atom spread between seeds; three-seed ensembles can.
  • A survivor by agreement. Three independent models vote, instead of one learned model extrapolating (E241).

Every contribution — the problem, the idea, how it was checked, what it bought, and whose work it builds on →

The ladderFive rungs, rising costWhat each rung asks, what it costs, how wrong it is, and what it cannot see.Open →The generatorA swarm on a connectomeThe walkers, the fly's whole brain as their head, and what the wiring does and does not add.Open →ExperimentsEvery entry in the log247 experiments, each with its question, prediction, result and verdict.Open →LogbookWhat changed, and whyNine eras: what was believed, built, measured, withdrawn and replaced.Open →RunsReplay recorded runsThe connectome's activity and the evaluator's answers, decision by decision.Open →BookThe method, in plain words14 chapters, each built around its diagrams: the question, the lattice, short-range order, when order appears, the cluster expansion and how it is learned, Monte Carlo, leaving the lattice, kinetics, DFT, the ladder, the generator, and what the ladder found.Open →

The generator

The MaleCNS rendering of the male Drosophila central nervous system
The MaleCNS v1.0 central nervous system, rendered by the dataset's authors (FlyEM / HHMI Janelia, University of Cambridge, MRC LMB, Google Research; male-cns.janelia.org, CC BY 4.0). The generator's head is this reconstruction's wiring: 164,506 neurons and 25.1 million connections (E139, E156).
generator diagram
The generator, as drawn for this site.

Twenty-four walkers move through the simplex of compositions. Each proposal is encoded as a sensory input — broadcast by a fixed random map onto all 17,479 afferent neurons of the MaleCNS connectome (164,506 neurons, 25.1 million connections) — propagated for four recurrent steps, and scored by a linear readout over the output neurons of the mushroom body, the lateral accessory lobe and the descending neurons. The readout is learned online from the rewards the ladder's cheapest rung pays; nearly all of its weight settles on the mushroom body's outputs (E161, E163).

What it has shown, on a reward that ranks: the readout learns the reward — score–reward correlation +0.90 against −0.03 for a frozen twin (E201) — and the swarm finds a quarter more distinct qualifying alloys than an elitist search at the same rate, 299 ± 31 against 238 ± 14 over three seeds (E195c). What it has not shown: its head-derived gain field was the worst of three tested (E198b), its plastic site in the fan-shaped body never reached the readout (N6), and the fixes meant to make it exploit harder made it worse (E202, E203). The advantage it has is width, not depth. Whether the connectome's wiring contributes anything a matched random graph does not is the question E224 was designed to settle.

E224 · in progress

Does the connectome's wiring matter? Eight arms, five seeds each, paired by seed: the real wiring (A), a degree-preserving rewire (B), a random head of matched sparsity (C), a frozen twin (D), and four baselines — elitist search, elitist with the fly's diversity devices, a ridge readout driving the same swarm, and MAP-Elites (E–H). Seeds recorded so far: A 5, B 5, C 5, D 4, E–H 0 of 20.

SeedReal wiring (A)Rewired (B)A − B
0333301+32
1278332−54
2333315+18
315240−225
4203334−131

Distinct qualifying finds over 800 verifications. The bar written before the run: all five A − B differences of one sign, with a median of at least 31. 2 of 5 differences are positive and the median is −54. On the five pairs recorded that bar is not met. Seed 3 was a cold-start stall for the real and the random head — qualifying finds per 50-round window 0 / 0 / 1 / 34 and 0 / 0 / 0 / 1 — which a replay is examining. No verdict is drawn until every arm is in.

Read at build time from runs/odt_map/stage_b_e224_<arm>_s<seed>.json; no runs/e224_card.txt yet.

The ladder

ladder diagram
The fidelity ladder, as drawn for this site.

Each rung asks a question the rung below cannot, and only survivors climb. The cheap rungs are the same model; the dear ones are independent of it.

Rung Asks Model Cost Measured error
0 · screen How far below its competitors is the random solid solution? Embedded cluster expansion (pyeCE), three-seed ensemble on RHEA DFT labels 0.1–0.5 s 7.5 meV/atom on eleven DFT cells no model chose (E237)
1 · ordering At what temperature does it order? Metropolis Monte Carlo on a 432-site bcc cell, 2600 → 60 K, the ensemble mean evaluated exactly 31 min median per find, rung 2 included (E240 walk logs) Mo–Ta 1149 ± 114 K; published 2020 ± 545 K (E242)
2 · lattice Is it bcc at all, once its atoms may move? MACE-MPA-0 relaxation against a 671-phase off-lattice hull 8–34 s 1–4 meV/atom on the gaps the ranking uses (E78)
3 · kinetics Can atoms travel far enough to reach what the thermodynamics wants? MACE vacancy formation and 30 migration barriers, gated on the 97.5th-percentile reach 52 min Recorded, not trusted, until an alloy anchor lands (E221)
4 · first principles Does density-functional theory agree? Quantum ESPRESSO, PBE, 60/720 Ry, 54-atom cells 3.4–4.5 h The reference the others are checked against

What is established

What is open

How it got here

Forager began as an idea around 1 September; the first files were made on 10 September and the first commit landed on 11 September. It started as a different experiment: a whole-brain network trained by a local forward-forward rule, choosing which of 64 fixed alloys to compute next. It did not choose — the three runs sampled uniformly with a fixed seed — and backpropagation reproduced its picks 39 times in 40 (E1), so the learning rule was not the problem. Forward-forward training was given up. Everything since, including what was withdrawn and why, is in the Logbook; the September runs are kept on the Results page.

Built with PRISMWebsite and visualizations made using Claude