Experiments · E183

Does the energy model fail across the hafnium–zirconium-rich region, not just at one cell?

Yes. A second cell came out +89.7 meV/atom in DFT against the model's −35.1; fitted reference energies for hafnium and zirconium were the cause.

In the log: a second cell at the edge: is the failure systematic?

confirmedDate 2026-09-18 09:12, as written in the logrung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 11646–11708, lines 11811–11836
exp E183 diagram
What E183 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E183.svg).

Pre-registration

The pre-registration, as written

E183 — a second cell at the edge: is the failure systematic?

If the eCE fails at one Hf/Zr-rich composition it may fail across the region, or E172 may be a single bad corner. Cell: Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ (Hf 0.26 Zr 0.24 Ti 0.15 Nb 0.20 Ta 0.15 — hcp-former-rich, no V, no Mo/W), seed-0 decoration, Vegard a from the same table, the settings standard, references refs_v3_std. v4 for this decoration, before DFT: −35.1 meV/atom (Vegard a = 3.4833 Å) — favourable, where E172's Hf/Zr-rich cell with Mo/V was +100 in DFT; v3b is not scored (retired). Predictions.

  1. |E_form(DFT) − v4| > 30 meV/atom, with DFT the more positive (v4 under-reads how unfavourable hcp-former-rich bcc mixtures are, as it did at E172). Then the edge failure is regional and the refit needs edge cells as a class, not E172 alone.
  2. If |DFT − v4| < 15, E172 was a corner and v4 covers the Hf/Zr region better than one cell suggested; the refit's first target is E172's neighbourhood only.
  3. Converges in < 40 iterations at the standard (five species, all with spn PAWs). Everything else in the comparison is clean: both sides QE, both at their own minima, alloy volume correction 0.02 meV, all references converged to 1e-9 Ry.

What the 46 means. This cell is Hf₁₀Zr₁₀Ti₂ with Mo₁₅V₉ — a hcp-former-rich random decoration, the corner of the nine-element space where RHEA's bcc data is thinnest and where the eCE's holdout error (a composition-wise average) says least. The eCE is not "wrong by 46 everywhere"; it is wrong by 46 here, on its first cell outside the data's centre, in the direction of under-predicting how unfavourable the mixture is. One cell is one cell: a second random decoration at a refractory-centred composition (MoNbTaW-like) is the obvious next rung-4 point, and its prediction goes in first.

Branch — the flywheel's first turn. This is exactly the case the design was built for: a rung-4 verdict the cheap rung did not see coming. The cell enters eCE training as the first QE-source point with a per-source point term (E164's branch), the eCE is refit, and the holdout error is re-measured; the prediction for that refit is that this cell's residual falls below 15 meV while the VASP holdout MAE moves by < 1 meV (one point among 4,300 cannot move the average). If the holdout MAE rises by more than 1, the QE point is fighting the VASP labels and the per-source term is not absorbing the code offset — that is the Δ-gauge question again, with a number.

2026-09-18 09:12 — the machine panicked, and I caused it. Kernel panic, "watchdog timeout: no checkins from watchdogd in 92 seconds", reboot. At the time: 8 pw.x ranks (E183, 60/720 Ry, 36 k-points) + three searches (E173, E173c, and the third-seed chain behind it) + the test suite, on 12 cores. RAM was not the limit (jobs ~6 GB of 26); the cores were oversubscribed until the system's own watchdog starved. Lost: E183 at SCF iteration 1 (nothing), E173 ~3 h into seed 0, E173c ~1.5 h, one suite run (its ".FFF.F" tail is thrash, not a regression — to be re-run in a free slot). Intact: every record, the DFT store (9 files), ece_v4_pureref, every finished verdict. No result recorded above depends on a job that died. Rule now in OVERNIGHT.md and in memory: pw.x ≤ 6 ranks, ONE search beside it, the suite only in an empty slot, everything nice -n 10, nothing launched above load 14. Also corrected: my hand-written trace times had drifted up to four hours from the clock; traces are stamped from date from here on (2026-09-18 09:16).

Colab CPU smoke (2026-09-18 09:44) — measured, and the answer is no for the searches as written. Operator: "use the colab cli". One CPU session (colab new): 2 cores Xeon 2.20 GHz, 12 GB, Python 3.13. Payload 100 MB (code + the six files a search reads); a single 100 MB upload broke the CLI's tunnel in under a second, 12 parts of 8 MB uploaded cleanly and reassembled to the same sha256. icet, ase, scipy installed on 3.13 without trouble; FORAGER_NO_CALC=1 (the searches never call MACE). E171b's protocol, 4 rounds × 4 × 1 seed = 16 verifications: not finished after 25 minutes with both cores saturated (load 2.8); the Mac does 16 verifications in about 3. The hot loop is a sparse mat-vec over 25.1 M synapses, memory-bandwidth-bound; the VM is ≥ 8× slower at it. One 800-verification seed would take ≥ 20 h — beyond the session limit and far beyond what the CLI's handle survives (1 h token, keep-alive). Session stopped; nothing left running on the account. What Colab is good for here: GPUs. The same mat-vec as a torch sparse CSR product on a T4/L4 should beat the M4, not trail it — that needs an optional device path in the propagation (_states, forward) with an equivalence test against scipy (scores equal to 1e-6 on CPU torch) before any GPU run. Also the eCE fits and, later, MACE fine-tuning. Not built tonight.

Results

EXPERIMENTS.md · line 11811

E183 result (13:02) — prediction 1 CONFIRMED, and the cause is not what E181 said. Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ at the standard, 22 iterations: E_form(DFT) = +89.7 meV/atom vs v4 −35.1 — wrong sign, 125 off. Two hcp-former-rich cells (E172: +100 vs +12; E183: +90 vs −35) make the edge failure regional. But it is not extrapolation: RHEA holds 554 training cells with Hf+Zr+Ti > 0.5, and their corrected labels average −47 meV/atom. The eCE reproduces its labels; the labels disagree with DFT by ~140 in that region. E181's "the model's own extrapolation failure" is withdrawn.

The mechanism, measured (references implied by e_lattice − e_formation, least squares):

el   ref used    lowest raw bcc pure   ref − raw     how the ref was set
Mo   −10.9154      −10.8503              −65 meV      corrected pure cell (correct: rattle removed)
Ta   −11.7940      −11.7490              −45          corrected pure cell
Hf   −12.3721      −12.5219             **+150**      fitted intercept
Ti    −7.6645       −7.6837              +19          fitted intercept
Zr    −8.1505       −7.9579             **−193**      fitted intercept

A corrected reference sits below the raw rattled minimum by the rattle energy (Mo, Ta: 45–65 meV). Hf's fitted intercept sits 150 meV above it — 200–300 meV above where a corrected bcc-Hf cell would be — so every Hf-rich formation energy is pushed favourable by ~0.2–0.3 × x_Hf eV; Zr's is 190 below, Ti's near neutral. The --pure-references pass fixed the six refractories from corrected pure cells and fitted the three hcp-formers; that choice is the defect. Branch → E186: corrected bcc pure references for Hf/Ti/Zr from RHEA's own bcc pure cells (16 / 6 / 3 exist), all nine fixed, refit → ece_v5. Prediction: v5 re-predicts E172 and E183 within 30 meV of +100 / +90, the refractory four stay within 0.8–1.3, and the holdout slope stays ≥ 0.8.

The full record

This entry is written in 2 separate places in the log, shown here in log order.

EXPERIMENTS.md · lines 11646–11708

E183 — a second cell at the edge: is the failure systematic?

If the eCE fails at one Hf/Zr-rich composition it may fail across the region, or E172 may be a single bad corner. Cell: Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ (Hf 0.26 Zr 0.24 Ti 0.15 Nb 0.20 Ta 0.15 — hcp-former-rich, no V, no Mo/W), seed-0 decoration, Vegard a from the same table, the settings standard, references refs_v3_std. v4 for this decoration, before DFT: −35.1 meV/atom (Vegard a = 3.4833 Å) — favourable, where E172's Hf/Zr-rich cell with Mo/V was +100 in DFT; v3b is not scored (retired). Predictions.

  1. |E_form(DFT) − v4| > 30 meV/atom, with DFT the more positive (v4 under-reads how unfavourable hcp-former-rich bcc mixtures are, as it did at E172). Then the edge failure is regional and the refit needs edge cells as a class, not E172 alone.
  2. If |DFT − v4| < 15, E172 was a corner and v4 covers the Hf/Zr region better than one cell suggested; the refit's first target is E172's neighbourhood only.
  3. Converges in < 40 iterations at the standard (five species, all with spn PAWs). Everything else in the comparison is clean: both sides QE, both at their own minima, alloy volume correction 0.02 meV, all references converged to 1e-9 Ry.

What the 46 means. This cell is Hf₁₀Zr₁₀Ti₂ with Mo₁₅V₉ — a hcp-former-rich random decoration, the corner of the nine-element space where RHEA's bcc data is thinnest and where the eCE's holdout error (a composition-wise average) says least. The eCE is not "wrong by 46 everywhere"; it is wrong by 46 here, on its first cell outside the data's centre, in the direction of under-predicting how unfavourable the mixture is. One cell is one cell: a second random decoration at a refractory-centred composition (MoNbTaW-like) is the obvious next rung-4 point, and its prediction goes in first.

Branch — the flywheel's first turn. This is exactly the case the design was built for: a rung-4 verdict the cheap rung did not see coming. The cell enters eCE training as the first QE-source point with a per-source point term (E164's branch), the eCE is refit, and the holdout error is re-measured; the prediction for that refit is that this cell's residual falls below 15 meV while the VASP holdout MAE moves by < 1 meV (one point among 4,300 cannot move the average). If the holdout MAE rises by more than 1, the QE point is fighting the VASP labels and the per-source term is not absorbing the code offset — that is the Δ-gauge question again, with a number.

2026-09-18 09:12 — the machine panicked, and I caused it. Kernel panic, "watchdog timeout: no checkins from watchdogd in 92 seconds", reboot. At the time: 8 pw.x ranks (E183, 60/720 Ry, 36 k-points) + three searches (E173, E173c, and the third-seed chain behind it) + the test suite, on 12 cores. RAM was not the limit (jobs ~6 GB of 26); the cores were oversubscribed until the system's own watchdog starved. Lost: E183 at SCF iteration 1 (nothing), E173 ~3 h into seed 0, E173c ~1.5 h, one suite run (its ".FFF.F" tail is thrash, not a regression — to be re-run in a free slot). Intact: every record, the DFT store (9 files), ece_v4_pureref, every finished verdict. No result recorded above depends on a job that died. Rule now in OVERNIGHT.md and in memory: pw.x ≤ 6 ranks, ONE search beside it, the suite only in an empty slot, everything nice -n 10, nothing launched above load 14. Also corrected: my hand-written trace times had drifted up to four hours from the clock; traces are stamped from date from here on (2026-09-18 09:16).

Colab CPU smoke (2026-09-18 09:44) — measured, and the answer is no for the searches as written. Operator: "use the colab cli". One CPU session (colab new): 2 cores Xeon 2.20 GHz, 12 GB, Python 3.13. Payload 100 MB (code + the six files a search reads); a single 100 MB upload broke the CLI's tunnel in under a second, 12 parts of 8 MB uploaded cleanly and reassembled to the same sha256. icet, ase, scipy installed on 3.13 without trouble; FORAGER_NO_CALC=1 (the searches never call MACE). E171b's protocol, 4 rounds × 4 × 1 seed = 16 verifications: not finished after 25 minutes with both cores saturated (load 2.8); the Mac does 16 verifications in about 3. The hot loop is a sparse mat-vec over 25.1 M synapses, memory-bandwidth-bound; the VM is ≥ 8× slower at it. One 800-verification seed would take ≥ 20 h — beyond the session limit and far beyond what the CLI's handle survives (1 h token, keep-alive). Session stopped; nothing left running on the account. What Colab is good for here: GPUs. The same mat-vec as a torch sparse CSR product on a T4/L4 should beat the M4, not trail it — that needs an optional device path in the propagation (_states, forward) with an equivalence test against scipy (scores equal to 1e-6 on CPU torch) before any GPU run. Also the eCE fits and, later, MACE fine-tuning. Not built tonight.

EXPERIMENTS.md · lines 11811–11836

E183 result (13:02) — prediction 1 CONFIRMED, and the cause is not what E181 said. Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ at the standard, 22 iterations: E_form(DFT) = +89.7 meV/atom vs v4 −35.1 — wrong sign, 125 off. Two hcp-former-rich cells (E172: +100 vs +12; E183: +90 vs −35) make the edge failure regional. But it is not extrapolation: RHEA holds 554 training cells with Hf+Zr+Ti > 0.5, and their corrected labels average −47 meV/atom. The eCE reproduces its labels; the labels disagree with DFT by ~140 in that region. E181's "the model's own extrapolation failure" is withdrawn.

The mechanism, measured (references implied by e_lattice − e_formation, least squares):

el   ref used    lowest raw bcc pure   ref − raw     how the ref was set
Mo   −10.9154      −10.8503              −65 meV      corrected pure cell (correct: rattle removed)
Ta   −11.7940      −11.7490              −45          corrected pure cell
Hf   −12.3721      −12.5219             **+150**      fitted intercept
Ti    −7.6645       −7.6837              +19          fitted intercept
Zr    −8.1505       −7.9579             **−193**      fitted intercept

A corrected reference sits below the raw rattled minimum by the rattle energy (Mo, Ta: 45–65 meV). Hf's fitted intercept sits 150 meV above it — 200–300 meV above where a corrected bcc-Hf cell would be — so every Hf-rich formation energy is pushed favourable by ~0.2–0.3 × x_Hf eV; Zr's is 190 below, Ti's near neutral. The --pure-references pass fixed the six refractories from corrected pure cells and fitted the three hcp-formers; that choice is the defect. Branch → E186: corrected bcc pure references for Hf/Ti/Zr from RHEA's own bcc pure cells (16 / 6 / 3 exist), all nine fixed, refit → ece_v5. Prediction: v5 re-predicts E172 and E183 within 30 meV of +100 / +90, the refractory four stay within 0.8–1.3, and the holdout slope stays ≥ 0.8.

Related entries

Built with PRISMWebsite and visualizations made using Claude