Experiment log · 2026-09-11 → 2026-09-24

Experiments

Every entry in the project's experiment log, from E1 to E242: what was asked, what was predicted before it ran, and what the log recorded afterwards — including the ones that failed, were withdrawn or were stopped.

Each entry has its own page with its pre-registration, its numbered predictions and the verdict written against each, every result paragraph, and the full text from EXPERIMENTS.md with line numbers. The Logbook tells the same history as a story.

252entries in the log: 247 experiments, E1–E242 with lettered variants, and 5 pre-registration documents
32state numbered predictions before the result
171have at least one result paragraph
14days, 2026-09-11 to 2026-09-24; 252 git commits

252 entries carry an illustration. 252 were checked against the files they name; 212 of those checks list discrepancies, shown on the entry page.

By status

A status is the catalog's reading of the log's own words: an explicit VOID, WITHDRAWN, superseded or STOPPED marker first, then the verdicts written against the numbered predictions. An entry whose text states no verdict word is recorded. The final check then re-read each entry against the files it names; where the keyword reading was wrong (150 entries, often an entry that withdraws an earlier claim), the checked status is shown and the entry page names the catalog's reading beside it.

By topic

Topic is the catalog's fixed keyword rule on the title and text (first match wins), so it is coarse: an entry about the fly that mentions DFT is filed under rung 4.

The log, played forward

dated in the log no date in the log: placed, in log order, inside the git window in which it was first written a git commit
Dates: the log's own date and time where it states them (59 entries to the minute or ten, 45 to the day, spread over that day in log order). For the 148 it does not date, the interval between the last commit of EXPERIMENTS.md without the entry and the first with it (the pre-registration files: the commit that added them). Sources: runs/site_data/experiments.json, git log. Click a dot to open its entry.

  1. E36Is the blended-element energy path truly smooth when computed exactly rather than sampled?Yes. Exactly, it is a cubic curve, and seven of eight element pairs change without a single turning point.unclassifiedillustratedcheck: 1 discrepancySep 24recorded1 result
  2. E225Does magnetism's energy depend on how the atoms are arranged, not just on the mix?Not yet. Paused: the first magnetic DFT cell never converged, oscillating for 180 steps.rung 4 · DFTillustratedcheck: 1 discrepancySep 24stopped1 result
  3. E242Does the chosen alloy really stay disordered when checked with DFT?Not yet. The DFT is still running; the model's own check on Mo–Ta passed, ordering at 1149 ± 114 K as expected.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 23running2 results
  4. E241Do independent models agree on one surviving find worth testing with DFT?Yes. Ten of 11 finds got all three votes; Mo₅₁Ti₃₈W₃Ta₃ was chosen, with a pass probability of 0.771.rung 4 · DFTillustratedcheckedSep 23recorded1 result
  5. E240Do the search's finds still pass when re-checked on the corrected ladder?Partly. After two scoring defects were fixed, 11 of 32 finds pass; the first pick was void, and Mo–Ta-rich finds did not order higher.rung 4 · DFTillustratedcheck: 1 discrepancySep 23mixed1 result
  6. E239Can a few dozen of our own DFT cells pin down each element's reference energy?Not yet. The 48-cell run was interrupted after 10 cells, restarted, and is still running; no correction has been measured.rung 4 · DFTillustratedcheck: 1 discrepancySep 23running1 result
  7. E238Does refitting the energy model on the larger new training set give a better model?No. Every new version over-bound the search's alloys (13.6–19.6 meV/atom error, bar 8), so the earlier model stays.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 23mixed3 results
  8. E237Can a mid-sized energy model get Mo–Ta right and still predict the search's alloys?Yes. Six numbers per element did both, with 6.1 meV/atom error on the search's alloys; it became the model that ships.rung 4 · DFTillustratedcheckedSep 23confirmed1 result
  9. E236Do the search's finds change when re-checked with the more accurate energy model?Not yet. That model was 55 times too slow for the ordering step; the run was stopped with nothing recorded and redone later.rung 4 · DFTillustratedcheckedSep 23stopped1 result
  10. E234Does adding the small ordered cells the labeller had skipped improve the energy model?Not yet. The cells were labelled (2,031 of 2,059), but the fit crashed on a data bug; they went into later fits instead.rung 1 · orderingillustratedcheck: 1 discrepancySep 22stopped0 results
  11. E233Does letting atom pairs interact over a longer distance make the energy model more accurate?Partly. Error on 11 DFT test cells fell from 20.6 to 15.9 meV/atom, just missing the bar of 15; the Mo–Ta miss did not shrink.rung 1 · orderingillustratedcheck: 2 discrepanciesSep 22mixed1 result
  12. E232Are the database's energies for strongly ordered cells right, checked against our own DFT?Partly. As used, all four sit within 13 meV/atom of DFT; the model, not the data, missed them. One corrected label (Cr–Ti) is 21.9 off.rung 4 · DFTillustratedcheckedSep 22mixed1 result
  13. E231Does recomputing every training energy at its true lowest point improve the energy model?Partly. Fits became steadier (seed spread 8.5 to 6.3 meV/atom), but the overall error barely moved and the Mo–Ta miss stayed.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 22mixed1 result
  14. E230Do the store's two DFT settings give the same energy for one cell?Yes. They differ by 4.4 meV/atom, under the 5 meV line (a gap over 15 was predicted), so the older rows stay.rung 4 · DFTillustratedcheck: 1 discrepancySep 22falsified1 result
  15. E229Would giving the energy model more numbers per element fix its Mo–Ta error?Partly. Error on 11 DFT test cells fell from 20.6 to 8.8 meV/atom, but unseen alloys got worse (7.1 to 11.0).rung 4 · DFTillustratedcheck: 1 discrepancySep 22mixed2 results
  16. E228Are the training energies biased because the machine-learned potential is too soft?No. The potential is not soft (force slope 1.014), but a volume-fitting step leaves labels 17.2 meV/atom too high.rung 4 · DFTillustratedcheckedSep 22mixed1 result
  17. E226Does switching on magnetism for Ni3Al keep the DFT reference frame consistent?Yes. The lattice constant held at 3.5726 A and the formation energy moved 40.7 meV/atom, inside the 100 limit.rung 4 · DFTillustratedcheck: 1 discrepancySep 22confirmed2 results
  18. E227Would a larger 54-atom cell confirm the magnetic-arrangement effect?Not yet. It was never launched; it waits on the magnetic-arrangement test, which is paused.rung 4 · DFTillustratedcheckedSep 22recorded1 result
  19. E224Does the real fly wiring find more alloys than the same brain randomly rewired?Not yet. Three of eight arms have finished and the result is not written; one seed stalled for 150 rounds.generator · fly brainillustratedcheck: 2 discrepanciesSep 22running1 result
  20. E223Can alloys proposed by a search pass the corrected ladder's ordering and stability checks?No. Of 32 finds, 24 could not pass because the simulation stopped at 100 K; none reached the 0.8 success bar.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 22superseded1 result
  21. E215cIs the sampler's temperature error exactly a factor of two?Yes. Measured on a fine 20 K grid, the factor is 2.000, to within 1.4 %.rung 1 · orderingillustratedcheck: 1 discrepancySep 22confirmed1 result
  22. E222Does chromium order against tantalum, as the energy model says, though a published model disagrees?Yes. DFT gives an ordering energy of 44.9 meV/atom, though the model's 107.9 is 2.4 times too big.rung 4 · DFTillustratedcheck: 1 discrepancySep 22confirmed2 results
  23. E221Does the machine-learned potential get vacancy energies right inside the MoNbTaW alloy?Not yet. Three of four DFT cells have finished; the tungsten cell failed and no result is written.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 21running0 results
  24. E190bWas the low ordering temperature a coarse-simulation artefact or a real model weakness?Withdrawn. The finer run found a peak at 367 K; doubled for the sampler's error it is 734 K, on the published 745 K.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 21withdrawn3 results
  25. E210bDoes the corrected sampler reproduce published ordering temperatures across nine alloys?Partly. Six comparable systems land within 22 % of the published values; Cr-Ta-Ti-W is off by a factor of 2.92.rung 4 · DFTillustratedcheck: 3 discrepanciesSep 21mixed0 results
  26. E220Does the energy model use real interactions beyond the second neighbour shell?No. Its third and fourth shells only echo what a nearest-neighbour model induces; the second shell is the one real extra coupling.rung 1 · orderingillustratedcheck: 1 discrepancySep 21mixed1 result
  27. E215Does the ordering sampler get the right temperature on a model with a known answer?No. It placed a known 1300-1410 K transition at 612 K, a factor of 2.1-2.3 too low.rung 4 · DFTillustratedcheck: 1 discrepancySep 21confirmed2 results
  28. E216Is ignoring magnetism a small error for chromium but large for iron, nickel, cobalt?Yes. Magnetism is worth 12 meV/atom in chromium but 61 to 467 meV/atom in nickel, cobalt and iron.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 21confirmed1 result
  29. E217Is the energy model reliable for a chromium-containing alloy when checked against DFT?Yes. For Cr-Ta-Ti-V-W it was within 15.1 meV/atom of DFT, inside the 25 meV/atom bar.rung 4 · DFTillustratedcheck: 1 discrepancySep 21confirmed1 result
  30. E218Does the DFT set-up reproduce known facts about Ni3Al, its first fcc alloy?Yes. Lattice constant 3.5706 A against experiment's 3.572, and ordering lowers the energy by 127.6 meV/atom.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 21confirmed2 results
  31. E219Does the machine-learned potential overestimate the energy of a vacancy in molybdenum?No. It matched DFT to 0.06 eV (3.107 against 3.164 eV), where an error above 0.3 eV was predicted.rung 4 · DFTillustratedcheck: 1 discrepancySep 21falsified1 result
  32. E214Was a higher-temperature ordering peak hidden above the simulation's old 1400 K ceiling?No. Swept up to 3000 K, the only peak stayed near 480 K; the ceiling was never the problem.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 21mixed2 results
  33. E213bDoes swapping two atoms in the ordered crystal cost what a nearest-neighbour model predicts?Yes. One far swap costs 9.96 meV/atom, 0.88 of the nearest-neighbour estimate and inside the predicted band.rung 1 · orderingillustratedcheck: 1 discrepancySep 21confirmed1 result
  34. E213Is the energy model's ordering energy mostly a nearest-neighbour effect?Yes. Nearest-neighbour order alone explains 99.3 % of the energy change along the ordering path; further shells add 0.3 %.rung 4 · DFTillustratedcheck: 1 discrepancySep 21mixed1 result
  35. E211Does the energy model find sensible lowest-energy arrangements, and does DFT agree with their energies?Partly. The arrangements are sensible, but DFT puts the model 32 meV/atom too shallow on Mo-Ta and 46.5 too deep on MoNbTaVW.rung 4 · DFTillustratedcheck: 1 discrepancySep 20mixed7 results
  36. E210Does the reversible ordering test give sensible temperatures on all nine published systems?Withdrawn. It stopped after its first system; an independent second sampler ran all nine systems instead.rung 1 · orderingillustratedcheck: 3 discrepanciesSep 20superseded0 results
  37. E209Does making walkers compete only within their own niche give deeper finds without losing numbers?Not yet. It was never run; like the five-element test, it waits on an energy-model refit that passes its tests.rung 1 · orderingillustratedcheck: 1 discrepancySep 20recorded1 result
  38. E208Can the search still find good alloys if each must mix five elements?Not yet. It was never run: it waits on an energy-model refit that passes its tests, and a trial run failed.generator · fly brainillustratedcheck: 2 discrepanciesSep 20recorded0 results
  39. E207Did the ordering simulation settle properly, or did it cool too fast?Yes. Heating and cooling agreed within 47 K; blaming the model for low temperatures was withdrawn once the sampler proved to run at double temperature.rung 1 · orderingillustratedcheck: 3 discrepanciesSep 20falsified2 results
  40. E196Do two more of the search's best finds hold up in full quantum calculations?Partly. One matched within 2.7 meV/atom; the molybdenum–niobium one came out 28 meV/atom shallower than the energy model said.rung 4 · DFTillustratedcheck: 1 discrepancySep 20mixed2 results
  41. E194Does the search's best new alloy hold up in a full quantum calculation?Yes. The quantum calculation gave −88.1 meV/atom against the energy model's −96.2, inside the 15 meV bar.rung 4 · DFTillustratedcheck: 1 discrepancySep 20confirmed1 result
  42. E198bWith the reward's gate removed, is the brain's own gain signal still the worst?Yes. It scored 136.9 against 153.2 for a fixed random gain, so the signal is better switched off.rung 1 · orderingillustratedcheck: 1 discrepancySep 19confirmed2 results
  43. E201bWith the reward's distorting gate removed, does the fly beat keep-the-best-and-mutate?Yes. The fly found 284 distinct alloys against 246, beyond the run-to-run spread, at a tied rate.generator · fly brainillustratedcheck: 1 discrepancySep 19mixed0 results
  44. E173cDoes the search do better with the serotonin learning site than without it?Partly. The score was about twice as high with the site (32.97 against 14.64), but the runs used different code, so it is a lead.unclassifiedillustratedcheck: 1 discrepancySep 19recorded1 result
  45. E202Does letting walkers search around their own finds make the fly find more?No. Finds halved, from 284 to 146.5, against a bar of 307; anchoring gave away the fly's wide coverage.generator · fly brainillustratedcheckedSep 19falsified1 result
  46. E203Does letting the brain imagine several candidates and pick the best give deeper finds?No. Finds got shallower (−62.9 against a bar of −72 meV/atom) and fewer (221 against 261).generator · fly brainillustratedcheck: 1 discrepancySep 19falsified1 result
  47. E204Does a reward-prediction-error learning rule make the brain's score track the reward better?Not yet. The run finished, but its score against the pre-registered bars was never written into the log.rung 1 · orderingillustratedcheck: 2 discrepanciesSep 19recorded1 result
  48. E205Do all three generator fixes together let the fly beat keep-the-best-and-mutate?No. The fly found 149 alloys against elitist's 238, and more slowly; two of the three fixes had already failed alone.generator · fly brainillustratedcheck: 2 discrepanciesSep 19falsified0 results
  49. E201Does the brain's readout actually learn the reward, compared with a frozen copy?Yes. The learning readout scored 78.0 against 15.0 for the frozen one, and its scores ranked the reward at 0.90.rung 0 · energy modelillustratedcheck: 2 discrepanciesSep 19mixed1 result
  50. E190Does the new energy model reproduce published ordering temperatures for nine alloys?Withdrawn. It scored 0 of 10 within 200 K, but the test was too coarse and its temperatures were half the true ones.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19falsified1 result
  51. E189Does the gain still help if each walker keeps following one other walker's live signal?Yes. On the old reward, freezing the pairing scored 41.1, level with the other live signals and 15 above a fixed random gain.unclassifiedillustratedcheck: 2 discrepanciesSep 16–19confirmed1 result
  52. E199Does teaching one brain region, once the signal can reach it, improve the search?Withdrawn. The run was declared void: its lessons were coin flips produced by a shared random-number fault.unclassifiedillustratedcheck: 2 discrepanciesSep 19void3 results
  53. E188At what temperature does half-molybdenum, half-tantalum order, according to the new energy model?Withdrawn. It read 500 ± 50 K, but the sampler ran at twice its stated temperature; on the corrected axis it is about 1000 K.rung 4 · DFTillustratedcheck: 3 discrepanciesSep 16–19superseded1 result
  54. E187Does the energy model hold up on the alloy where two cheap models disagree most?Yes. The quantum calculation gave −76.4 meV/atom against the energy model's −78.7; the older model was off by 65.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19confirmed1 result
  55. E197Does searching on the refitted energy model give deeper, more varied finds?No. The finds' spread fell to about 15 meV/atom against a bar of 40, and the best find got shallower, not deeper.rung 1 · orderingillustratedcheck: 1 discrepancySep 19mixed0 results
  56. E186Did fixing the pure elements' reference energies make the energy model match quantum calculations?Yes. All six quantum-calculated cells now agree within about 10 meV/atom; only a side check, the holdout slope (0.746 against 0.8), missed.rung 4 · DFTillustratedcheck: 1 discrepancySep 16–19mixed2 results
  57. E195cWith a third run each, does the fly still only tie the best simple search?No. The fly found 299 distinct alloys against 238, well beyond the run-to-run spread, at the same rate.generator · fly brainillustratedcheck: 1 discrepancySep 19falsified2 results
  58. E195bWith foreign elements removed, does the fly brain beat simple searches that use no brain?Withdrawn. It tied keep-the-best-and-mutate on finds (158 against 169.5) and lost on rate, but this reversed once the reward's gate was removed.generator · fly brainillustratedcheck: 2 discrepanciesSep 19superseded4 results
  59. E200Does a bigger simulation box raise molybdenum–tantalum's ordering temperature?No. The larger box gave 457 ± 147 K against 500 ± 50 K; both are on the sampler's axis, half the true temperature.rung 1 · orderingillustratedcheck: 3 discrepanciesSep 19mixed0 results
  60. E191Does the energy model get the energy released by ordering right, checked by quantum calculation?Partly. For molybdenum–tantalum yes, 80.3 against 80.8 meV/atom; for the five-element alloy the ordered cell was a guess, so no ratio stands.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 19mixed0 results
  61. E185Is a gain that follows a live brain signal better than a fixed random one?Yes. On the old reward the fixed gain scored 25.5 against about 40 for live signals; on the real reward this later reversed.unclassifiedillustratedcheckedSep 16–19falsified1 result
  62. E198On the real reward, does the brain's own gain signal beat a fixed random one?No. The brain's own signal came last of three, 16.6 points below a fixed random gain, against a predicted lead of 15.generator · fly brainillustratedcheck: 2 discrepanciesSep 18falsified1 result
  63. E195On the new reward, does the fly brain beat simple searches that use no brain?Withdrawn. Keep-the-best-and-mutate matched or beat the fly here, but that reading reversed once a distorting gate was removed from the reward.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 18superseded1 result
  64. E184bDoes the search give the same result on the graphics chip as on the processor?Partly. One run matched within 0.29 points; the other missed the 1.0 bar by 0.67, so each comparison now stays on one device.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 16–19falsified3 results
  65. E172bDoes the model's 46 meV/atom miss survive stricter DFT settings?Yes. At stricter settings the cell came out at +99.8 meV/atom, only 2.6 from before, so the miss is the model's.unclassifiedillustratedcheckedSep 16–19confirmed1 result
  66. E173Once the confirmation step is switched on, does the serotonin site actually get taught?Yes. It received 138 and 187 lessons in the two seeds; whether that helps the search was left to a control run.generator · fly brainillustratedcheck: 1 discrepancySep 16–19recorded1 result
  67. E181Does retraining the energy model on labels referenced to pure elements fix its DFT misses?Partly. Four refractory cells now match DFT to about 10 %, but the hafnium–zirconium-rich cell got worse: +12 against +97 meV/atom.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19mixed1 result
  68. E192Are the ordered states found by the model's own simulation real, checked by quantum calculation?Partly. For the five-element alloy the model is 25 meV/atom too deep on every such state; for molybdenum–tantalum it is 23 too shallow.rung 4 · DFTillustratedcheck: 3 discrepanciesSep 18mixed4 results
  69. E193Once the search is paid by the new energy model, does it rank its finds?Yes. Finds now spread over 60 meV/atom instead of all scoring about −11; one depth bar was narrowly missed (68% and 79% against 80%).rung 4 · DFTillustratedcheck: 2 discrepanciesSep 18mixed1 result
  70. E180Do the energy model's training labels get the molybdenum–tantalum mixing energy right?Partly. DFT gave −89.8 meV/atom: near the labels referenced to pure elements (−73.7), far from the fitted-reference labels the model used (−22.9).rung 4 · DFTillustratedcheckedSep 16–19recorded1 result
  71. E179Does a different random arrangement of the same composition give the same DFT energy?Yes. The second arrangement gave −72.5 meV/atom against −71.1, a 1.4 difference, well inside the 12.4 allowance.rung 4 · DFTillustratedcheck: 1 discrepancySep 16–19confirmed1 result
  72. E178Did the navigation-centre learning site, rather than the extra propagation step, lift the score?No. At two seeds the run without the site (34.77) fell between both bars; a third seed later matched the site's run exactly.unclassifiedillustratedcheck: 2 discrepanciesSep 16–19recorded1 result
  73. E177At a composition the search itself found, does DFT agree with the cheap models?No. DFT gave −71.1 meV/atom against −17.9 and −10.9 from the two models; most of the gap was later traced to the labels' references.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19falsified1 result
  74. E184Could the brain's slowest calculation run on the Mac's graphics chip, and how much faster?Yes. One walker move fell from 1.88 s to 0.04 s on the graphics chip, 47 times faster; learning stayed on the main processor.unclassifiedillustratedcheck: 2 discrepanciesSep 18recorded0 results
  75. E171bDoes switching on the navigation centre's learning site make the whole-brain search better?Withdrawn. The score reached 39.96, but a later run without the site matched it exactly; the extra propagation step did the work.generator · fly brainillustratedcheck: 2 discrepanciesSep 16–19withdrawn2 results
  76. E183Does the energy model fail across the hafnium–zirconium-rich region, not just at one cell?Yes. A second cell came out +89.7 meV/atom in DFT against the model's −35.1; fitted reference energies for hafnium and zirconium were the cause.rung 4 · DFTillustratedcheck: 3 discrepanciesSep 18confirmed1 result
  77. E176At the centre of its training data, does the cheap model agree with DFT?Yes. DFT gave −49.1 meV/atom against −21.9, inside the allowed band; the factor-two gap was later traced to the labels' reference energies.rung 4 · DFTillustratedcheck: 1 discrepancySep 16–19confirmed1 result
  78. E172Does the cheap energy model agree with full DFT on a first 54-atom test cell?No. DFT gave +97.2 meV/atom against the model's +51.2, a 46 meV/atom miss, well outside the allowed band.rung 4 · DFTillustratedcheck: 1 discrepancySep 16–19falsified2 results
  79. E170Must each walker's stride come from the brain's reading at that walker's own spot?No. Handing each walker another walker's reading scored as well (31.5 against 25.9); the spread of strides helped, not the brain.generator · fly brainillustratedcheck: 2 discrepanciesSep 16–19falsified1 result
  80. E171Does letting the input spread more steps wake up the brain's navigation centre?No. A fourth step reached every navigation-centre cell, but their activity stayed about 20 times below other regions and did not rise.rung 4 · DFTillustratedcheck: 3 discrepanciesSep 16–19falsified1 result
  81. E169Can the brain's steering cells choose the walkers' headings instead of a random draw?No. Steering from those cells sent the walkers into the same spot: one find per seed where the random draw had found about 71.unclassifiedillustratedcheck: 1 discrepancySep 16–19withdrawn1 result
  82. E167Does letting the brain set each walker's stride and patience improve the search?Withdrawn. The score more than doubled, to 25.9, but a shuffled control matched it: varied strides helped, not the brain's reading.generator · fly brainillustratedcheck: 2 discrepanciesSep 16–19withdrawn2 results
  83. E166Does a second learning site in the brain's navigation centre change the search?No. The run matched the one without it to every digit; the site's updates never reached anything the readout sees.rung 1 · orderingillustratedcheck: 2 discrepanciesSep 16–19falsified1 result
  84. E165Is the brain's own learning rule the one that makes the walkers do worse?Partly. It lowered the mushroom body's score, but on the whole brain it raised it by 0.057; there the readout's rule did the harm.rung 0 · energy modelillustratedcheckedSep 16–19mixed1 result
  85. E164Do the two DFT codes agree closely enough to mix their data without a correction?Partly. Energy differences agreed to 0.8 meV/atom for TaW but differed by 6.8 for VW, so a per-code correction term was adopted.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19mixed1 result
  86. E163Does feeding the input to every sensory cell make the whole-brain search better?Yes. The search score rose from 2.61 to 11.57, but its two seeds scored about 2.7 and 20, so the gain's size is uncertain.generator · fly brainillustratedcheck: 2 discrepanciesSep 16–19confirmed1 result
  87. E162Measured by the ladder's own verdict, do lessons move the walkers toward better alloys?No. With lessons on, the swarm's average ladder score fell on both brains, from 0.153 to 0.101 on the mushroom body.rung 0 · energy modelillustratedcheckedSep 16–19falsified1 result
  88. E161Did the learned readout put real weight on the newly read brain regions?No. The mushroom-body outputs kept 99.6 % of the weight and the new regions 0.4 %, so the earlier gain came from learning, not location.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19falsified3 results
  89. E160With lessons switched on, do the whole-brain walkers climb higher?No. Their climb fell by 0.58 standard deviations with lessons on, and the record warns that a learning readout stretches its own ruler.rung 0 · energy modelillustratedcheckedSep 16–19falsified1 result
  90. E159Does reading more of the fly brain, with learned weights, improve the search?Yes. The search score rose from 0.15 to 2.61; a later check found the gain came from learning the weights, not from the extra cells.rung 0 · energy modelillustratedcheck: 2 discrepanciesSep 16–19falsified1 result
  91. E158Does the new energy model rank atom arrangements better than the old one?Partly. It ranked better on unseen compositions (0.80 against 0.65) and on held-out MoNbTaVW, but the old model won on MoNbTaW (0.83 against 0.52).rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16–19mixed2 results
  92. E157Does feeding every sensory cell wake more of the fly brain than smell alone?Yes. Every sensory cell woke 25.1 % of the brain and reached the steering cell; smell alone woke at most 4.2 %.rung 1 · orderingillustratedcheck: 2 discrepanciesSep 16–19mixed2 results
  93. E156Does the whole brain contain the fly's own steering circuit, reachable from the mushroom body?Yes. Mushroom-body outputs reach the steering centres directly (9.2 % land in one of them), but the model reads none of these paths.generator · fly brainillustratedcheck: 1 discrepancySep 16–19recorded0 results
  94. E155Do the walkers climb worse on the whole brain's landscape than on the mushroom body's?Yes. They relocate more and climb less; learning adds 0.27 standard deviations of climb on the mushroom body but only 0.03 on the whole brain.generator · fly brainillustratedcheckedSep 16–19mixed1 result
  95. E154Do the extra neurons reach the score through paths that learning cannot change?Partly. They do, but even the mushroom body gets 76 % of its readout through unlearnable paths; the whole brain costs ×2–3, not ×10.generator · fly brainillustratedcheck: 1 discrepancySep 16–19mixed3 results
  96. E153Do the extra neurons weaken the brain's composition code before learning happens?No. Input gain was identical and Kenyon-cell input grew only 3 %; a one-third weaker drive is absorbed by the threshold.generator · fly brainillustratedcheck: 2 discrepanciesSep 16–19falsified2 results
  97. E152With everything else matched, does the whole brain search worse than the mushroom body alone?Yes. The mushroom body found 5.5 alloys per run, the whole brain 0.5, although both read the same 4,064 Kenyon cells.generator · fly brainillustratedcheck: 1 discrepancySep 16–19superseded4 results
  98. E151Are the corrected labels real signal, or mostly the correction's own error?Partly. Below 0.10 Å of strain they are usable: noise 12.4 meV/atom against a 29 meV/atom ordering signal.rung 4 · DFTillustratedcheckedSep 16–19mixed7 results
  99. E150Can RHEA's rattled and squeezed energies be corrected into clean training labels?Partly. Formation energies then centred at 0.0 meV/atom instead of +622, but spread wider than predicted until strained frames were cut.rung 1 · orderingillustratedcheck: 1 discrepancySep 16–19falsified2 results
  100. E149Was the new energy model being trained on structures that can show atoms ordering?No. The first training set held 3,758 random structures and none of the 4,750 ordered ones; both runs were stopped.rung 1 · orderingillustratedcheck: 1 discrepancySep 16–19stopped0 results
  101. E147Can a database of known crystal structures tell whether a found alloy is new?No. Its 671 phases are compounds with at most two elements; the nearest to any find sat 0.263 away, beyond the 0.15 tolerance.rung 2 · hull, MACEillustratedcheckedSep 16confirmed1 result
  102. E146Do three safety checks written into the code actually do anything?No. A magnetic warning was dropped, a key guard was never used and two settings were ignored; two are now wired, one still is not.rung 4 · DFTillustratedcheckedSep 16recorded1 result
  103. E145With nothing else changed, does the whole fly brain search better than the mushroom body?No. It found 0.5 alloys per run against the mushroom body's 5.5 in the matched control, and its best was −33 meV/atom.generator · fly brainillustratedcheck: 2 discrepanciesSep 16falsified0 results
  104. E144Did the rediscovery discount itself cause the collapse in finds?Yes. One number both paid the fly and counted finds, so a good alloy near a known one scored 0.240, under the 0.832 bar.unclassifiedillustratedcheckedSep 16recorded0 results
  105. E143Does the ordering scan fail when a transition lies near the top of the scan?Yes. MoNbTaTiW's two runs disagreed by 396 K, and the fit reached past its 2,914 K melting point; the alloy still passes.rung 1 · orderingillustratedcheck: 1 discrepancySep 16recorded0 results
  106. E142Is the quick ordering gauge reliable for equal-share alloys, as earlier assumed?No. It under-read all three, at 0.09 to 0.77 of the simulated value, so no single correction factor rescues it.rung 1 · orderingillustratedcheck: 3 discrepanciesSep 16confirmed1 result
  107. E141Is rung 1's documented run-to-run scatter of 23 K right?No. Five runs of one alloy scatter by 46 K, twice the documented 23 K; the claim that scan range never matters was later withdrawn.rung 1 · orderingillustratedcheck: 1 discrepancySep 16withdrawn1 result
  108. E140Does paying the fly less for rediscoveries push it to find more different alloys?Withdrawn. A bug let the discount erase finds (5 fell to 1), and the run changed two things at once, so it settles nothing.generator · fly brainillustratedcheck: 1 discrepancySep 16void0 results
  109. E139Does running the generator on the whole fly brain, through every check, find more alloys?Not yet. No result was written for this run; the whole-brain question was answered later, one change at a time.unclassifiedillustratedcheck: 2 discrepanciesSep 16superseded0 results
  110. E138Does the ordering temperature depend on how hot the simulation's scan goes?Partly. Not when the alloy orders well inside the scan; near or beyond the scan's top, the answer moved by 972 K.rung 1 · orderingillustratedcheck: 1 discrepancySep 16mixed0 results
  111. E137Was the generator using the fly's whole brain?No. It used 5,311 neurons — the mushroom body and its inputs — about 2.5 % of the 211,577 annotated cells.generator · fly brainillustratedcheck: 2 discrepanciesSep 16recorded0 results
  112. E136Does the fly-brain generator match a standard search, and is anything it finds new?Partly. Its best was 4 meV/atom short of the standard search's, but it found 5 alloys to 34.5; 'new' meant only 'unlisted'.generator · fly brainillustratedcheck: 2 discrepanciesSep 16mixed0 results
  113. E135Do the two best five-element refractory alloys really order inside the 90–1000 K window?Partly. MoNbTaVW orders inside, near 690 K; MoNbTaTiW stays above 1,565 K, though later checks showed its exact temperature is not known.rung 1 · orderingillustratedcheck: 2 discrepanciesSep 16falsified2 results
  114. E134Are five-element refractory alloys shut out by physics, or only by how they were sampled?Withdrawn. The claim that physics excludes them fell: a free search reached −82 meV/atom, past the −40 bar that the grid's best (−30.3) missed.generator · fly brainillustratedcheck: 3 discrepanciesSep 16withdrawn1 result
  115. E133Does a free search find passing alloys that a fixed grid of 1,705 points missed?Partly. 11 off-grid compositions passed against the grid's 2, but grid point Mo–Ta 50/50 stayed best, and the 'copper family' was just Mo–Ta.search baselinesillustratedcheck: 3 discrepanciesSep 16mixed1 result
  116. E132Does removing the discredited 'too slow to happen' credit shrink the feasible set further?Yes. It fell from 12 to 2 (Mo–Ta and Ta–W), not the 4 predicted; only equal-share Mo–Ta is cleared by both checks.rung 3 · kineticsillustratedcheck: 1 discrepancySep 16mixed2 results
  117. E131Do the 55 'feasible' alloys survive once two discredited checks stop being trusted?No. 43 of the 55 rested on an ordering reading of exactly 0 K, which is no clearance; 12 remained, and only equal-share Mo–Ta robustly.rung 1 · orderingillustratedcheck: 1 discrepancySep 16confirmed2 results
  118. E130Is the cheap energy model much worse against quantum calculations than its advertised error?No. With lattice size and atom positions matched, its error is 8.1 meV/atom, below the advertised 9.09; the earlier 44.6 was a mismatch.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16mixed2 results
  119. E129Was the 44.6 meV/atom error a comparison of like with like?No. The model was scored on perfect lattice sites against quantum energies of displaced atoms, so relaxation energy counted as error.rung 4 · DFTillustratedcheck: 1 discrepancySep 16confirmed1 result
  120. E128Is the energy model far less accurate against quantum calculations than its fit suggests?Withdrawn. It seemed to be 44.6 meV/atom, but that counted relaxation energy as error; a matched comparison later gave about 8.rung 4 · DFTillustratedcheck: 1 discrepancySep 16withdrawn1 result
  121. E127With the heat-capacity reading repaired, do our ordering temperatures now match the published scale?Not yet. The rerun was set up but no result was ever recorded; the predicted ratio was 0.65 to 0.80, up from 0.54.rung 1 · orderingillustratedcheck: 1 discrepancySep 16stopped0 results
  122. E126Did the larger set of competing phases, not the new energy model, shift old results?No. The added phases moved refractory results by at most 0.8 meV/atom; the new energy model moved them by 4.5 meV/atom typically.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 16mixed1 result
  123. E125Given a budget big enough to start, does the diversity-seeking search find more qualifying alloys?No. At 1200 evaluations it found 2.67 qualifying alloys per run against 54 for the standard search, though its best reached −76 meV/atom.search baselinesillustratedcheck: 1 discrepancySep 16mixed1 result
  124. E124Should the diversity-seeking search become the default at the usual budget of 240 evaluations?No. It samples at random for its first 600 evaluations, so at 240 it never starts; its best was +129 against −48 meV/atom.search baselinesillustratedcheck: 1 discrepancySep 16confirmed1 result
  125. E123Does reporting the best find from each niche, instead of the top five, reveal more?Yes. It surfaced 18 compositions instead of 5 at no cost to the best find, but only one of the new ones qualified.unclassifiedillustratedcheck: 1 discrepancySep 16confirmed1 result
  126. E122Can the ladder flag alloys holding elements with their own transition in the working range?Yes. It flags 10 of the 55 candidates, and those hold the only elements here that resist oxygen well; the 55 later shrank to 12.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 16mixed1 result
  127. E121Can we recover how the default twelve-element energy model was built?Yes. The build logs were on disk: 2961 structures and a fit error of 9.09 meV/atom, not the 8.07 our spec quoted.unclassifiedillustratedcheck: 1 discrepancySep 16recorded0 results
  128. E120Is the log's count of 613 competing phases simply out of date?Yes. The file holds 671: the same 136 prototypes plus 535 database phases. But 18 data files have no recorded producer.rung 2 · hull, MACEillustratedcheckedSep 16mixed1 result
  129. Pre-reg · acquisitionCould the fly's mushroom body choose which alloys deserve an expensive calculation?Not yet. The test was designed and its two pass gates fixed in advance, but it was never run.generator · fly brainillustratedcheck: 1 discrepancySep 16recorded0 results
  130. Pre-reg · arenaCan the fly's brain steer through alloy space with fewer calculations than simple controllers?Not yet. The arena and its controls were written and the kill rule fixed in advance, but the test was never run.generator · fly brainillustratedcheck: 1 discrepancySep 16recorded0 results
  131. Pre-reg · wholebrainDoes the real fly wiring search alloy space better than a shuffled copy of itself?No. The real wiring lost to a shuffle (−244.4 vs −259.0 meV/atom, lower is better), and a 7 × 8 table did as well.unclassifiedillustratedcheck: 1 discrepancySep 16falsified0 results
  132. E119Written down rung by rung, does the checking ladder do what we thought it did?No. The cheapest check is not in the ladder's walk, the second runs at its worst setting, and two rungs were never timed.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  133. E118Among the stable candidates, which balance oxidation, weight and brittleness best?Withdrawn. A 14-member trade-off front favoured chromium–nickel, but its candidate list was later cut from 55 to 12.unclassifiedillustratedcheck: 1 discrepancySep 13–16withdrawn0 results
  134. E117Does rewarding low energy steer the search toward alloys that fail the ordering test?No. The opposite: top-rewarded compositions pass both tests 12 times more often than the grid overall; later corrections raised this to 34.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16falsified1 result
  135. E116On twelve elements, does the search still collapse onto one corner of the composition space?Yes. It collapses at least as hard, onto vanadium and tungsten; the statistical flaw first blamed was then measured and ruled out.search baselinesillustratedcheck: 1 discrepancySep 13–16mixed2 results
  136. E115With its set-up defects fixed, does the diversity-seeking search match the best single-goal search?Yes. Its best find came within 0.5 meV/atom, it covered 4.8 times more of the map, and it found a second stable family.search baselinesillustratedcheck: 2 discrepanciesSep 13–16confirmed1 result
  137. E114Was the search exploring all twelve elements, and does a diversity-seeking search do better?Withdrawn. Only eight elements were switched on; the diversity search's 98.8 meV/atom loss came from set-up defects, and fixed it matched the best.search baselinesillustratedcheck: 1 discrepancySep 13–16falsified1 result
  138. E113Does running more independent copies of the sampler remove its error in ordering temperature?Partly. More copies shrink the scatter, but only longer runs remove the bias: 11 per cent high at 90 sweeps, 1 per cent at 1000.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16recorded0 results
  139. E112Do different pairs of elements drive the two ordering transitions, as published?Partly. Nb–W orders at the colder transition (333 K) and Mo–Ta at the warmer (667 K), but their ranges overlap.rung 1 · orderingillustratedcheckedSep 13–16mixed1 result
  140. E111Is the published reference we compared against itself biased high?Yes. It is a mean-field estimate that runs about 1.3 times high; with our baseline fixed too, the factor of two shrinks to 1.04.unclassifiedillustratedcheckedSep 13–16recorded0 results
  141. E110Does the way the heat capacity is read place the transition too cold?Partly. The flat baseline biased it cold (571 K became 770 K once fixed), but there were two transitions, at 333 and 733 K.rung 1 · orderingillustratedcheckedSep 13–16mixed1 result
  142. E109Is the energy model to blame for ordering temperatures coming out half the published values?No. It overestimates ordering energies by 4 to 65 per cent on five alloys, so the fault lies in the sampling.rung 1 · orderingillustratedcheckedSep 13–16falsified1 result
  143. E108Do our predicted ordering temperatures agree with 75 published ones?Partly. They rank alloys only weakly (correlation 0.44) and come out at 0.54 of the published scale on 74 of 75 alloys.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16mixed1 result
  144. E107Would a newer learned energy model that shares information between elements serve us better?Yes. It probably should replace ours: it stays compact as elements are added and reached 4 meV/atom on six elements.rung 0 · energy modelillustratedcheck: 1 discrepancySep 13–16recorded0 results
  145. E106Which known defects can give wrong answers, and can the worst be fixed first?Partly. Eleven defects were ranked by harm and the two worst were fixed: a broken DFT comparison and a backwards kinetics charge.unclassifiedillustratedcheckedSep 13–16recorded0 results
  146. E105Was the checking ladder giving up on alloys too early, at its second step?Yes. With the early-exit rule reversed the ladder reaches the top: MoNbTaW passes with probability 0.94 instead of 0.00006.unclassifiedillustratedcheckedSep 13–16confirmed1 result
  147. E104Would the screen's claim that chromium–nickel is stable survive a proper relaxation?No. Relaxed properly, chromium–nickel sits 86 meV/atom above the stability floor; the uncertainty term was right to distrust it.unclassifiedillustratedcheck: 2 discrepanciesSep 13–16confirmed0 results
  148. E103Did the generator really never find the chromium–nickel alloy?No. It found Cr0.62Ni0.38, ranked 32 of 74; only the top eight rows were ever reported, so the earlier claim is withdrawn.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded1 result
  149. E102Is there a published search method that keeps every good family of alloys, not just one?Yes. MAP-Elites (2015) keeps the best alloy in each cell of a feature grid and has found crystal structures, but it was not run here.generator · fly brainillustratedcheck: 1 discrepancySep 13–16recorded0 results
  150. E101Is there a high-energy ridge between the tungsten–tantalum and chromium–nickel alloys?Yes. The ridge is 403 meV/atom high, but it was later shown not to explain the convergence; the reward's uncertainty penalty did.rung 3 · kineticsillustratedcheck: 1 discrepancySep 13–16superseded1 result
  151. E100Does the generator return only tantalum–tungsten alloys because the other elements are genuinely worse?Withdrawn. Other elements do score well (Cr0.64Ni0.36 at −82 meV/atom), but the generator had found Cr–Ni; the report showed only the top eight.rung 3 · kineticsillustratedcheck: 1 discrepancySep 13–16withdrawn1 result
  152. E99Does the cheap ordering check miss orderings in alloys away from a 50/50 mix?Yes. For W0.73 Ta0.24 it reported no ordering, while the simulation found one at 649 K, inside the window.rung 1 · orderingillustratedcheckedSep 13–16confirmed1 result
  153. E98Can the energy model itself estimate each alloy's ordering temperature in a few milliseconds?Withdrawn. It ran in 1–4 ms, but it misses orderings away from 50/50 mixes; a 1.2-second short simulation now clears those.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16mixed1 result
  154. E97Can the ordering temperature be estimated cheaply from known ordered structures in the databases?No. It underestimated by 38–72 % because every database ordering is binary, and it mis-ranked two of the four alloys.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16falsified1 result
  155. E96Was the screen quietly rewarding alloys that would order inside the temperature window?Yes. Three of the four chosen alloys order at 844–991 K, inside the window, so the best-binary ranking was withdrawn.rung 3 · kineticsillustratedcheck: 1 discrepancySep 13–16confirmed1 result
  156. E95Do the alloys the screen lets through cluster in melting point?No. Only 2.8 % pass, and they cluster by few elements (3 ± 1) and small size mismatch (2 ± 1 %), not by melting point.unclassifiedillustratedcheckedSep 13–16mixed1 result
  157. E94Can a cheap activation-energy formula be fitted on 35 alloys spanning two to twelve elements?Not yet. The survey stopped after six alloys; two nickel-rich ones would not stay bcc, and five of six were alloys the screen rejects anyway.unclassifiedillustratedcheck: 1 discrepancySep 13–16stopped1 result
  158. E93Is the estimate's error explained by the spread of hopping barriers in mixed alloys?Partly. Easy paths do dominate, but mostly through vacancy formation, not the barrier spread (0.265 to 0.321 eV); no cheap estimate was validated.unclassifiedillustratedcheck: 1 discrepancySep 13–16falsified1 result
  159. E92Does the melting-point estimate of the activation energy hold for alloys with many elements?No. It overestimates by up to 1.58 eV on an eight-element alloy, the direction that fakes stability, so the previous result was withdrawn.unclassifiedillustratedcheckedSep 13–16falsified1 result
  160. E91If the screen credits alloys whose atoms are too slow to move, does the search diversify?Withdrawn. Diversity rose (34 % of finds had five or more elements), but it came from an overestimated activation energy; the term was taken out.rung 3 · kineticsillustratedcheck: 2 discrepanciesSep 13–16withdrawn1 result
  161. E90Would rewarding exploration make the generator propose alloys with more elements?No. The optimistic reward saw one more element but found fewer alloys and none with four or more; the screen itself favours simple binaries.rung 3 · kineticsillustratedcheck: 1 discrepancySep 13–16falsified2 results
  162. E89Does adding training data where the search goes remove the model's excess caution there?Yes. 1,535 extra structures halved Mo–Ta's uncertainty (63 to 32 meV/atom); its top ranking was later withdrawn for a different reason.unclassifiedillustratedcheck: 1 discrepancySep 13–16confirmed1 result
  163. E88Given chromium, cobalt, copper and nickel too, does the generator still pick refractory metals?Yes. All ten top picks were tantalum–tungsten alloys with none of the new elements; the best score, −75 meV/atom, was just outside the predicted range.generator · fly brainillustratedcheck: 1 discrepancySep 13–16mixed1 result
  164. E87Is the coarse sampling grid used in the DFT calculations accurate enough?Yes. The series converges (6×6×6 and 8×8×8 agree to 0.17 meV/atom), and the 3×3×3 grid in use costs about 3.8 meV/atom.rung 4 · DFTillustratedcheckedSep 13–16recorded1 result
  165. E86Can counting each atom's neighbours separate bcc-like structures from close-packed ones?Yes. Bcc-like structures count 8 neighbours and close-packed ones 12; 401 bcc orderings were set aside once two faults in the filter were fixed.rung 3 · kineticsillustratedcheck: 1 discrepancySep 13–16confirmed1 result
  166. E85Can one energy model cover twelve elements without losing accuracy on the original eight?Yes. The error rose only to 9.09 meV/atom, but a disguised bcc ordering in the database hull flipped every refractory verdict, so that hull was withdrawn.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 13–16confirmed1 result
  167. E84Does taking competing phases from materials databases catch what hand-picked structures missed?Withdrawn. It fixed the nickel compounds, but the hull also held bcc orderings, so it was replaced by a filtered hull.unclassifiedillustratedcheck: 1 discrepancySep 13–16superseded0 results
  168. E83Can one energy model cover nickel as well as the eight refractory metals?Withdrawn. The fit mixed two reference conventions 226 meV/atom apart; a corrected refit on the shared references reached a 7.41 meV/atom error.unclassifiedillustratedcheck: 3 discrepanciesSep 13–16withdrawn1 result
  169. E82Do the four earlier verdicts survive the refitted energy model?Yes. All four still pass; the two controls reproduced exactly, and the ordering temperatures moved less than predicted, by −37 to +133 K.unclassifiedillustratedcheck: 1 discrepancySep 13–16mixed0 results
  170. E81Does the list of competing phases already cover what nickel forms with these metals?No. Nickel compounds sit 63–236 meV/atom below the solid solution, and three of the guessed structures were simply the wrong ones.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 13–16falsified1 result
  171. E80Can the energy model be made exact for pure elements without losing accuracy elsewhere?Yes. Weighting the eight pure elements tenfold cut their worst error from 57.5 to 4 meV/atom; accuracy elsewhere barely moved, but ordering temperatures shifted.rung 3 · kineticsillustratedcheckedSep 13–16confirmed1 result
  172. E79Once the model's vanadium error stops paying, what does the site's generator run find?Partly. Vanadium-rich picks vanished and the family stayed Mo–Nb–Ta–W, but the best score was −90 meV/atom, short of the predicted −110 to −120.generator · fly brainillustratedcheck: 2 discrepanciesSep 13–16mixed1 result
  173. E78Does first-principles physics rank the alloys the same way the fast potential does?Yes. The fast potential over-binds each alloy by about 30 meV/atom, but the differences between alloys agree to within 4 meV/atom.unclassifiedillustratedcheck: 1 discrepancySep 13–16confirmed0 results
  174. E77Can two DFT cells holding different elements be compared by subtracting their energies?No. Each element adds its own huge constant, giving gaps like 629,552 meV/atom; comparing mixing energies against each element's own reference fixes it.rung 4 · DFTillustratedcheck: 1 discrepancySep 13–16recorded1 result
  175. E76Could a bookkeeping bug make DFT seem to contradict the atomistic model?Yes. Energies were divided by the atom count twice, shrinking a real 100 meV/atom gap to 6; it was caught before any result was used.rung 4 · DFTillustratedcheck: 1 discrepancySep 13–16recorded1 result
  176. E75Will first-principles DFT agree with the atomistic model on which alloys are more stable?Yes. Tested later on a repaired comparison, the two agreed to 28 meV/atom on average, inside the predicted 30.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 13–16confirmed1 result
  177. E74Do the generator's own best compositions beat MoNbTaW through every check?Partly. All four are 1.9 to 2.3 times more stable against decomposition, but order at higher temperatures; they pass only because atoms move slowly.rung 2 · hull, MACEillustratedcheck: 2 discrepanciesSep 13–16superseded0 results
  178. E73Was the first DFT crash a problem with the calculation settings?No. Atoms sat 0.433 Å apart instead of 2.852 Å, because fractional positions were read as ångströms. A geometry check now refuses such cells.rung 4 · DFTillustratedcheckedSep 13–16recorded0 results
  179. E72Was the relaxation model to blame when pure vanadium screened as a stable alloy?No. The energy model itself was wrong at the pure elements, by up to 58 meV/atom at molybdenum; an uncertainty penalty now rejects them.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  180. E71Does the fly's learning advantage survive when its reward is corrected?Yes. It still beat its frozen copy (p = 0.0006), but the advantage shrank from 3.9 to 1.7 times.generator · fly brainillustratedcheck: 1 discrepancySep 13–16confirmed1 result
  181. E70Did the quick check's relaxation estimate use the right measure of strain?No. The whole cell's strain, not the atoms' size mismatch, predicts relaxation; the refitted model's error fell from 82 to 24 meV/atom.unclassifiedillustratedcheck: 3 discrepanciesSep 13–16falsified0 results
  182. E69Does the quick check's relaxation estimate work where the generator actually searches?No. There it predicts about 10–12 meV/atom everywhere while the truth varies by 38; it is off by −24 on average.unclassifiedillustratedcheck: 1 discrepancySep 13–16mixed0 results
  183. E68Are the quoted ordering-temperature error bars as wide as the model's own uncertainty?No. Model uncertainty dominates: the bars were 1.7 to 2.1 times too narrow (Mo–Ta: 177 K, not 86 K).rung 1 · orderingillustratedcheck: 3 discrepanciesSep 13–16mixed0 results
  184. E67Do proposals from a second, independent generator run also pass every check?Partly. All eight passed and a Mo–Ta binary was found twice, but ranking it the best alloy was later withdrawn.unclassifiedillustratedcheck: 3 discrepanciesSep 13–16superseded0 results
  185. E66Does running the atomistic model in single precision change any results?No. Forces agree to 0.02 meV/Å and relaxations take identical steps, while running 1.5 times faster.unclassifiedillustratedcheck: 1 discrepancySep 13–16confirmed0 results
  186. E65Is the energy model badly wrong outside the compositions it was trained on?Withdrawn. The comparison mixed conventions; done like-for-like, the model's error is 3 meV/atom, better than its stated 5.67.rung 2 · hull, MACEillustratedcheckedSep 13–16withdrawn1 result
  187. E64Does adding tungsten change stability, ordering temperature and atom movement?Partly. More tungsten steadily costs stability against decomposition (r = +0.94); its apparent effects on ordering temperature and migration barriers were confounding.unclassifiedillustratedcheck: 1 discrepancySep 13–16mixed0 results
  188. E63Can the generators be compared on the full requirement instead of the quick check?No. On equilibrium alone nearly every composition orders inside 90–1000 K, so every generator scores zero or near zero; kinetics must decide.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  189. E62Does tungsten keep these alloys stable by slowing atom movement?Withdrawn. The next experiment found tungsten makes vacancies rarer, not slower; the five-alloy trend was confounded. The candidates still passed.unclassifiedillustratedcheck: 2 discrepanciesSep 13–16superseded0 results
  190. E61Does the fly circuit actually learn once its starved coding layer is fixed?Yes. The learning fly then beat its frozen copy on all 16 seeds, though not the simple hill-climber.generator · fly brainillustratedcheck: 1 discrepancySep 13–16recorded0 results
  191. E60On the real target, does the fly-brain generator beat simpler generators?No. A simple hill-climber scored 1.8 times higher, and the fly did not beat its own frozen copy (p = 0.40).unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  192. E59Can a fast stability check scan the whole design space, and where does it point?Yes. At 4.4 ms per alloy, 20,000 random compositions all point to the Mo–Nb–Ta–W corner; only 0.41% were stable.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  193. E58Once atom movement is counted, does any alloy meet the requirement?Yes. MoNbTaW, which the survey rejected: its atoms move less than a lattice spacing in 1000 hours at 1000 K (pass probability 0.88).rung 3 · kineticsillustratedcheckedSep 13–16recorded0 results
  194. E57Can the energy to create a vacancy in these alloys be computed reliably?Partly. A formula error had made the spread ±1.36 eV (really ±0.04), but the absolute value still disagreed with the literature by about 1 eV.rung 0 · energy modelillustratedcheck: 2 discrepanciesSep 13–16recorded0 results
  195. E56Do the survey's qualifying alloys survive a check against every competing crystal structure?No. All seven sit 98 to 185 meV/atom above the competing phases at 90 K; the rejected MoNbTaW is the only one below.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  196. E55Did the fix to the pure-element reference energies actually reach the model in use?No. The model in use was never rebuilt; corrected at load, the equiatomic mixing energy moves from −152 to +73 meV/atom.unclassifiedillustratedcheck: 2 discrepanciesSep 13–16recorded0 results
  197. E54Are the survey's best alloys really stable, or does another crystal structure beat them?No. A Laves crystal structure, invisible to our lattice-only model, sits 119 to 123 meV/atom lower in energy at HfV2 and ZrV2.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  198. E53Was the 2.2-fold gap to published ordering temperatures a real physics problem?No. It came from the comparison table, the model slightly flattening order, and the estimator. Separately, every mixing energy was about 226 meV/atom too negative.unclassifiedillustratedcheck: 3 discrepanciesSep 13–16recorded0 results
  199. E52Was the ordering temperature truly that uncertain, or was our reading of it too noisy?No. Reading the centre of the whole heat-capacity curve, not its tallest point, cut the scatter nearly fivefold, from 271 K to 57 K.unclassifiedillustratedcheck: 2 discrepanciesSep 13–16recorded0 results
  200. E51Was the fly's learning rule correct, and were the ordering labels good enough to learn?No. The rule's sign was inverted (fixed, 0.665 became 0.875), and the labels were too noisy to learn from.labels · dataillustratedcheck: 2 discrepanciesSep 13–16recorded0 results
  201. E50Does the full screen, promote and learn loop beat random search?No. The loop runs and the fly learns (−0.295 to +0.580), but it only matches random search: its code has no notion of similar alloys.generator · fly brainillustratedcheck: 1 discrepancySep 13–16recorded0 results
  202. E49Tuned properly, does the whole brain steer the search better than a small matrix?No. It matched a 7 × 8 matrix to within 1.26 meV/atom and lost to random directions by 36 meV/atom; two pre-registered failure conditions fired.rung 1 · orderingillustratedcheck: 1 discrepancySep 13–16falsified0 results
  203. E48Can turning up the input make the brain model use its on-off switching?No. Switches set at zero scale exactly with the input; only thresholds solved neuron by neuron raised switching to 94.6% of neurons.generator · fly brainillustratedcheckedSep 13–16recorded0 results
  204. E47Does the whole fly brain compute anything a small table of 56 numbers cannot?No. At this setting a 7 × 8 matrix reproduced its steering almost exactly (similarity 0.980), and it searched 41 meV/atom worse than random directions.generator · fly brainillustratedcheck: 1 discrepancySep 13–16falsified0 results
  205. E46Does the fly's real brain wiring carry alloy information better than randomly rewired copies?No. Reading the composition off the output neurons scored 0.461 for the real wiring against 0.781 for shuffled copies.unclassifiedillustratedcheck: 1 discrepancySep 13–16recorded0 results
  206. E45Does a map of 219 alloys reproduce the published trends in ordering temperature?Partly. Titanium lowers it, as published, but weakly (rank correlation −0.149); niobium lowers it most (−0.478); temperatures span 150–1225 K.unclassifiedillustratedcheck: 1 discrepancySep 13recorded0 results
  207. E44Are the fly's proposals worse than evenly spread ones when only the proposer changes?No. The same optimiser reached −479.9 from the fly's proposals and −479.0 meV/atom from evenly spread ones: a draw, 1.2 standard errors.unclassifiedillustratedcheckedSep 13recorded0 results
  208. E43Can the quantum calculations be run in parallel in the cloud, cheaply?Yes. A 16-atom calculation took 1,592 s on four cores for $0.09, and each machine deleted itself when done.unclassifiedillustratedcheckedSep 13recorded0 results
  209. E42Does our chain reproduce the published ordering temperature of tantalum-titanium-vanadium-tungsten?Partly. We get 383 ± 75 K against their 500 K, about 120 K low, and the same strongest pair, tantalum with tungsten.unclassifiedillustratedcheck: 2 discrepanciesSep 12recorded1 result
  210. E41Does more training data pin down the ordering temperature?Withdrawn. It scattered 96–150 K however much data was used; that scatter was later traced to how the peak was read, not to the data.rung 1 · orderingillustratedcheck: 1 discrepancySep 12superseded0 results
  211. E40Can the fly be given honest error bars even though its error is bias?Yes. A calibration that assumes nothing reached 90% coverage; the fly's own code gave the tightest bars, 16% narrower than knowing nothing.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  212. E39Does our chain reproduce a published ordering temperature and the most strongly paired metals?Withdrawn. One run gave 700 K and the published strongest pair; repeated runs moved the pair and loosened the temperature to a bound.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 12superseded1 result
  213. E38Is one carefully arranged cell more accurate than one random arrangement of atoms?Yes. It misses the true average by 1.92 meV/atom against 5.54, 2.9 times better; a second potential disagreed with a spread of 8.8.rung 2 · hull, MACEillustratedcheckedSep 12recorded0 results
  214. E37Can the fly generate good new alloys that no pre-made list contained?No. It generated alloys outside every list but gained 0.0123 of the score against the optimiser's 0.0750, a pre-registered loss.generator · fly brainillustratedcheck: 1 discrepancySep 12falsified1 result
  215. Pre-reg · generativeCan the fly generate new alloys that beat ranking a fixed list, on two goals?No. Ranking a list with a standard statistical model gained 0.0750; the generating fly gained 0.0123, a pre-registered loss.unclassifiedillustratedcheck: 1 discrepancySep 12falsified0 results
  216. E35With two goals at once, does the fly beat the standard statistical optimiser?No. In a pre-registered test the optimiser won by 5.03 standard errors and the fly fell below random; the benchmark also had little power.generator · fly brainillustratedcheck: 3 discrepanciesSep 12falsified1 result
  217. Pre-reg · multiobjectiveDoes the fly beat a standard statistical model at choosing alloys for two goals?No. The standard model scored 16.93, random 16.45 and the fly 15.73: a pre-registered loss, below random.generator · fly brainillustratedcheck: 1 discrepancySep 12falsified0 results
  218. E34Does blending one element into another smooth the energy landscape, even for very different metals?Withdrawn. Replaced by an exact calculation: the answer is still yes, seven of eight paths are smooth, and the bumps seen here were sampling noise.unclassifiedillustratedcheck: 2 discrepanciesSep 12superseded1 result
  219. E33Does the fly find better alloys than the standard statistical optimiser people actually use?Withdrawn. The fly lost by 11.3 meV/atom, but that test mixed who proposes with who picks; a fairer test later found a draw.generator · fly brainillustratedcheck: 1 discrepancySep 12superseded0 results
  220. E32Can the fly decide for itself how long to keep searching before moving on?Yes. A stop-when-returns-fall rule matched the hand-set schedule (−470.6 against −471.6 meV/atom) with a tighter spread, though a minimum still binds.generator · fly brainillustratedcheck: 1 discrepancySep 12recorded1 result
  221. E31Was the reference simulation cell big enough for the energy model's reach?No. Its 4.94 Å limit was below the model's 6.0 Å reach; in a larger cell the accuracy figure is 1.33, not 0.96 meV/atom.unclassifiedillustratedcheckedSep 12recorded0 results
  222. E16bIs the eight-element energy formula accurate enough to build the rest of the work on?Yes. Its cross-checked error is 5.67 meV/atom, smaller than the 6.61 meV/atom error of the model it copies.rung 0 · energy modelillustratedcheck: 1 discrepancySep 12recorded1 result
  223. E30Does the fly, searching freely, find better alloys than random sampling at the same cost?Partly. It reached −471.6 against −456.8 meV/atom (t = 4.6), but a later random baseline on other seeds put that win in doubt.generator · fly brainillustratedcheck: 1 discrepancySep 12mixed1 result
  224. E29Can any of the fly's own signals tell where its predictions are wrong?No. Its error is 99.8% repeatable bias, so all three spread-based signals correlated negatively with it (−0.08 to −0.21).generator · fly brainillustratedcheck: 2 discrepanciesSep 12recorded0 results
  225. E28Does the fly rank alloys better when dopamine signals surprise rather than raw reward?Yes. Teaching from outcome minus expectation raised its held-out rank correlation to 0.853, against 0.561 for the composition alone.generator · fly brainillustratedcheck: 3 discrepanciesSep 12recorded0 results
  226. E27Could the fly's odour circuit learn to rank alloys it had not seen?Partly. Once its code was made sparse it learned (rank correlation −0.23 to +0.15), but only about a quarter of a straight-line fit's +0.56.generator · fly brainillustratedcheck: 3 discrepanciesSep 12recorded0 results
  227. E26Could the energy model be fitted directly to published quantum-mechanical data instead of a potential?No. That data sits off the lattice or in tiny two-element cells; fits to it err by 18–25 meV/atom, against 5.67 via the potential.rung 4 · DFTillustratedcheck: 1 discrepancySep 12recorded0 results
  228. E25Can the free energy steer the alloy search, and does it say anything new?Yes. It runs at 100–125 alloys per second and shows low free energy comes from ordering, a trade-off (correlation +0.670).unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  229. E24Does the spread of an alloy's random-arrangement energies predict how much it orders?Yes. It correlates at +0.877, against −0.21 for the old measure, and cuts the shortcut's error to 0.96 meV/atom.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  230. E23Can the slow free-energy calculation be replaced by a quick calibrated shortcut?Withdrawn. It matched to 1.60 meV/atom, but was calibrated on the wrong variable; a correction based on energy scatter replaced it.unclassifiedillustratedcheck: 1 discrepancySep 12superseded1 result
  231. E22Can existing public quantum datasets support a move to nickel superalloys?No. Not one public dataset holds a nickel–aluminium–chromium–cobalt structure; that data would have to be generated.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 12recorded1 result
  232. E21Is the usual assumption of perfectly random mixing good enough for the entropy?No. At 1500 K the real entropy is 3.4% lower, about 6 meV/atom, and at 300 K it is only half.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  233. E20Can a model that also predicts magnetism replace the main energy model?No. It is four times worse on energy (26.50 meV/atom), but it spots magnetic systems, so it serves as a warning flag.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 12recorded1 result
  234. E19Is the screening model equally reliable for every kind of atomic structure?No. Errors range from 6.51 to 119.50 meV/atom; intermetallic compounds, the key competing phases, are a weak spot at 39.09.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 12recorded0 results
  235. E18Is the new screening model accurate on independent quantum data for these alloys?Yes. Near normal volume its error is 6.61 meV/atom on 300 independent structures; squeezed or stretched cells are about ten times worse.rung 4 · DFTillustratedcheck: 2 discrepanciesSep 12recorded1 result
  236. E17Does the quantum calculation confirm that small simulation boxes are biased?No. The 16-to-54-atom shift is +4.65 ± 2.88 meV/atom, within noise. The screening model overstates the scatter about twofold.rung 4 · DFTillustratedcheckedSep 12recorded1 result
  237. E16Can a fast formula over neighbouring atoms stand in for the slow energy model?Yes. For four elements it matches the model to 0.76 meV/atom on unseen arrangements, at almost no cost.rung 0 · energy modelillustratedcheck: 1 discrepancySep 12recorded1 result
  238. E15Does learning at the fly's learning site close the gap to the simple formula?No. Learning moved the ranking by 0.01 against a gap of 0.37, although the weights changed a lot.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  239. E14Had the old screening model ever seen an alloy like the ones it was judging?No. Its training data holds not one metal alloy of four or more of these elements; a public set of 2,150 exists.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 12recorded2 results
  240. E13Would a newer version of the screening model fix its distorted energies?Yes. A newer, freely licensed model cut the error from 70.4 to 23.0 meV/atom and ranked the ten alloys best.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  241. E12Does the fly brain beat a simple straight-line formula on the eight element fractions?No. The formula reaches 0.948 rank agreement with the target; the untrained 166,700-neuron brain reaches 0.755.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  242. E11Does the real wiring rank alloys closer to the target than rewired copies do?Yes. The real wiring scores −0.755 against −0.679 for rewired copies, though a simple formula later did far better.rung 4 · DFTillustratedcheck: 1 discrepancySep 12recorded1 result
  243. E10Does the brain's actual wiring at the learning site change how it ranks alloys?Yes. Randomly rewiring those connections drops agreement with the original ranking to 0.868, far beyond seed noise (0.9995).unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  244. E9Can taking the wiring and the output signs from the fly's anatomy make the ranking stable?Yes. With each output neuron signed by its own teaching neurons, rankings from different seeds agree at 0.9995.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  245. E8Does reading a named group of output neurons give a stable ranking of alloys?Partly. It removed a size artefact, but different random seeds still ranked the alloys differently, correlations from −0.76 to +0.61.unclassifiedillustratedcheck: 1 discrepancySep 12recorded1 result
  246. E7Is the fly's own learning circuit present, complete, in the imported brain?Yes. Every pathway is there, with 61,210 connections at the learning site, and smell receptors correctly never connect to it directly.unclassifiedillustratedcheck: 2 discrepanciesSep 12recorded0 results
  247. E6Was the old importer keeping all of the fly brain's connections?No. Its five-contact cut-off discarded 76% of them; the full brain has 25,582,938 connections, extracted in 11.5 seconds.generator · fly brainillustratedcheck: 2 discrepanciesSep 12recorded1 result
  248. E5Can a 128-atom quantum calculation run on this computer?No. It needs more than 23.45 GB of memory against 15.6 GB available, and was killed after about 35 minutes.rung 4 · DFTillustratedcheck: 1 discrepancySep 12recorded1 result
  249. E4Does the quantum calculation also give different energies for different random arrangements?Yes. Three arrangements of one cell spread by 4 meV/atom; the book's −58 meV/atom lies inside that spread.rung 4 · DFTillustratedcheckedSep 12recorded0 results
  250. E3Must the simulated box of atoms be large before its energy can be trusted?Withdrawn. The screening model showed small boxes off by up to 15 meV/atom, but a later quantum check found no clear size effect.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 12superseded0 results
  251. E2Is one random arrangement of the atoms enough to rank alloys by energy?No. One arrangement varies by about 9 meV/atom, as much as the gaps between close rivals; only the coarse ranking is safe.rung 2 · hull, MACEillustratedcheck: 1 discrepancySep 12recorded1 result
  252. E1Was the learning rule the reason the brain failed to choose, or would backpropagation do better?No. Backpropagation reproduced the local rule's picks at 39 of 40 steps, so the fault lay upstream of the rule.unclassifiedillustratedcheck: 2 discrepanciesSep 11recorded1 result

Dates in grey italics are not stated in the log: they are the git window in which the entry was first written (hover for the exact commits). A day filter matches an undated entry if its window covers that day.

Built with PRISMWebsite and visualizations made using Claude