Does the energy model fail across the hafnium–zirconium-rich region, not just at one cell?
Yes. A second cell came out +89.7 meV/atom in DFT against the model's −35.1; fitted reference energies for hafnium and zirconium were the cause.
In the log: a second cell at the edge: is the failure systematic?
confirmedDate 2026-09-18 09:12, as written in the logrung 4 · DFT0 predictions · 1 result paragraphEXPERIMENTS.md lines 11646–11708, lines 11811–11836
What E183 did and how it came out, drawn from this record and the files it names (book/assets/diagrams/exp/E183.svg).
Pre-registration
The pre-registration, as written
E183 — a second cell at the edge: is the failure systematic?
If the eCE fails at one Hf/Zr-rich composition it may fail across the region, or E172 may
be a single bad corner. Cell: Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ (Hf 0.26 Zr 0.24 Ti 0.15 Nb 0.20 Ta 0.15
— hcp-former-rich, no V, no Mo/W), seed-0 decoration, Vegard a from the same table, the
settings standard, references refs_v3_std. v4 for this decoration, before DFT:
−35.1 meV/atom (Vegard a = 3.4833 Å) — favourable, where E172's Hf/Zr-rich cell with Mo/V
was +100 in DFT; v3b is not scored (retired).
Predictions.
|E_form(DFT) − v4| > 30 meV/atom, with DFT the more positive (v4 under-reads how
unfavourable hcp-former-rich bcc mixtures are, as it did at E172). Then the edge failure
is regional and the refit needs edge cells as a class, not E172 alone.
If |DFT − v4| < 15, E172 was a corner and v4 covers the Hf/Zr region better than one cell
suggested; the refit's first target is E172's neighbourhood only.
Converges in < 40 iterations at the standard (five species, all with spn PAWs). Everything else in the comparison is clean: both sides QE, both at their own
minima, alloy volume correction 0.02 meV, all references converged to 1e-9 Ry.
What the 46 means. This cell is Hf₁₀Zr₁₀Ti₂ with Mo₁₅V₉ — a hcp-former-rich random
decoration, the corner of the nine-element space where RHEA's bcc data is thinnest and
where the eCE's holdout error (a composition-wise average) says least. The eCE is not
"wrong by 46 everywhere"; it is wrong by 46 here, on its first cell outside the data's
centre, in the direction of under-predicting how unfavourable the mixture is. One cell is
one cell: a second random decoration at a refractory-centred composition (MoNbTaW-like)
is the obvious next rung-4 point, and its prediction goes in first.
Branch — the flywheel's first turn. This is exactly the case the design was built for:
a rung-4 verdict the cheap rung did not see coming. The cell enters eCE training as the
first QE-source point with a per-source point term (E164's branch), the eCE is refit, and
the holdout error is re-measured; the prediction for that refit is that this cell's
residual falls below 15 meV while the VASP holdout MAE moves by < 1 meV (one point among
4,300 cannot move the average). If the holdout MAE rises by more than 1, the QE point is
fighting the VASP labels and the per-source term is not absorbing the code offset — that
is the Δ-gauge question again, with a number.
2026-09-18 09:12 — the machine panicked, and I caused it. Kernel panic, "watchdog
timeout: no checkins from watchdogd in 92 seconds", reboot. At the time: 8 pw.x ranks
(E183, 60/720 Ry, 36 k-points) + three searches (E173, E173c, and the third-seed chain
behind it) + the test suite, on 12 cores. RAM was not the limit (jobs ~6 GB of 26); the
cores were oversubscribed until the system's own watchdog starved. Lost: E183 at SCF
iteration 1 (nothing), E173 ~3 h into seed 0, E173c ~1.5 h, one suite run (its ".FFF.F"
tail is thrash, not a regression — to be re-run in a free slot). Intact: every record,
the DFT store (9 files), ece_v4_pureref, every finished verdict. No result recorded above
depends on a job that died. Rule now in OVERNIGHT.md and in memory: pw.x ≤ 6 ranks,
ONE search beside it, the suite only in an empty slot, everything nice -n 10, nothing
launched above load 14. Also corrected: my hand-written trace times had drifted up to four
hours from the clock; traces are stamped from date from here on (2026-09-18 09:16).
Colab CPU smoke (2026-09-18 09:44) — measured, and the answer is no for the searches as written.
Operator: "use the colab cli". One CPU session (colab new): 2 cores Xeon 2.20 GHz, 12 GB,
Python 3.13. Payload 100 MB (code + the six files a search reads); a single 100 MB upload
broke the CLI's tunnel in under a second, 12 parts of 8 MB uploaded cleanly and
reassembled to the same sha256. icet, ase, scipy installed on 3.13 without trouble;
FORAGER_NO_CALC=1 (the searches never call MACE). E171b's protocol, 4 rounds × 4 × 1 seed =
16 verifications: not finished after 25 minutes with both cores saturated (load 2.8);
the Mac does 16 verifications in about 3. The hot loop is a sparse mat-vec over 25.1 M
synapses, memory-bandwidth-bound; the VM is ≥ 8× slower at it. One 800-verification seed
would take ≥ 20 h — beyond the session limit and far beyond what the CLI's handle survives
(1 h token, keep-alive). Session stopped; nothing left running on the account.
What Colab is good for here: GPUs. The same mat-vec as a torch sparse CSR product on a
T4/L4 should beat the M4, not trail it — that needs an optional device path in the
propagation (_states, forward) with an equivalence test against scipy (scores equal
to 1e-6 on CPU torch) before any GPU run. Also the eCE fits and, later, MACE fine-tuning.
Not built tonight.
Results
EXPERIMENTS.md · line 11811
E183 result (13:02) — prediction 1 CONFIRMED, and the cause is not what E181 said.
Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ at the standard, 22 iterations: E_form(DFT) = +89.7 meV/atom vs v4 −35.1
— wrong sign, 125 off. Two hcp-former-rich cells (E172: +100 vs +12; E183: +90 vs −35)
make the edge failure regional. But it is not extrapolation: RHEA holds 554 training
cells with Hf+Zr+Ti > 0.5, and their corrected labels average −47 meV/atom. The eCE
reproduces its labels; the labels disagree with DFT by ~140 in that region. E181's
"the model's own extrapolation failure" is withdrawn.
The mechanism, measured (references implied by e_lattice − e_formation, least squares):
el ref used lowest raw bcc pure ref − raw how the ref was set
Mo −10.9154 −10.8503 −65 meV corrected pure cell (correct: rattle removed)
Ta −11.7940 −11.7490 −45 corrected pure cell
Hf −12.3721 −12.5219 **+150** fitted intercept
Ti −7.6645 −7.6837 +19 fitted intercept
Zr −8.1505 −7.9579 **−193** fitted intercept
A corrected reference sits below the raw rattled minimum by the rattle energy (Mo, Ta:
45–65 meV). Hf's fitted intercept sits 150 meV above it — 200–300 meV above where a
corrected bcc-Hf cell would be — so every Hf-rich formation energy is pushed favourable by
~0.2–0.3 × x_Hf eV; Zr's is 190 below, Ti's near neutral. The --pure-references pass
fixed the six refractories from corrected pure cells and fitted the three hcp-formers;
that choice is the defect. Branch → E186: corrected bcc pure references for Hf/Ti/Zr
from RHEA's own bcc pure cells (16 / 6 / 3 exist), all nine fixed, refit → ece_v5.
Prediction: v5 re-predicts E172 and E183 within 30 meV of +100 / +90, the refractory four
stay within 0.8–1.3, and the holdout slope stays ≥ 0.8.
The full record
This entry is written in 2 separate places in the log, shown here in log order.
EXPERIMENTS.md · lines 11646–11708
E183 — a second cell at the edge: is the failure systematic?
If the eCE fails at one Hf/Zr-rich composition it may fail across the region, or E172 may
be a single bad corner. Cell: Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ (Hf 0.26 Zr 0.24 Ti 0.15 Nb 0.20 Ta 0.15
— hcp-former-rich, no V, no Mo/W), seed-0 decoration, Vegard a from the same table, the
settings standard, references refs_v3_std. v4 for this decoration, before DFT:
−35.1 meV/atom (Vegard a = 3.4833 Å) — favourable, where E172's Hf/Zr-rich cell with Mo/V
was +100 in DFT; v3b is not scored (retired).
Predictions.
|E_form(DFT) − v4| > 30 meV/atom, with DFT the more positive (v4 under-reads how
unfavourable hcp-former-rich bcc mixtures are, as it did at E172). Then the edge failure
is regional and the refit needs edge cells as a class, not E172 alone.
If |DFT − v4| < 15, E172 was a corner and v4 covers the Hf/Zr region better than one cell
suggested; the refit's first target is E172's neighbourhood only.
Converges in < 40 iterations at the standard (five species, all with spn PAWs). Everything else in the comparison is clean: both sides QE, both at their own
minima, alloy volume correction 0.02 meV, all references converged to 1e-9 Ry.
What the 46 means. This cell is Hf₁₀Zr₁₀Ti₂ with Mo₁₅V₉ — a hcp-former-rich random
decoration, the corner of the nine-element space where RHEA's bcc data is thinnest and
where the eCE's holdout error (a composition-wise average) says least. The eCE is not
"wrong by 46 everywhere"; it is wrong by 46 here, on its first cell outside the data's
centre, in the direction of under-predicting how unfavourable the mixture is. One cell is
one cell: a second random decoration at a refractory-centred composition (MoNbTaW-like)
is the obvious next rung-4 point, and its prediction goes in first.
Branch — the flywheel's first turn. This is exactly the case the design was built for:
a rung-4 verdict the cheap rung did not see coming. The cell enters eCE training as the
first QE-source point with a per-source point term (E164's branch), the eCE is refit, and
the holdout error is re-measured; the prediction for that refit is that this cell's
residual falls below 15 meV while the VASP holdout MAE moves by < 1 meV (one point among
4,300 cannot move the average). If the holdout MAE rises by more than 1, the QE point is
fighting the VASP labels and the per-source term is not absorbing the code offset — that
is the Δ-gauge question again, with a number.
2026-09-18 09:12 — the machine panicked, and I caused it. Kernel panic, "watchdog
timeout: no checkins from watchdogd in 92 seconds", reboot. At the time: 8 pw.x ranks
(E183, 60/720 Ry, 36 k-points) + three searches (E173, E173c, and the third-seed chain
behind it) + the test suite, on 12 cores. RAM was not the limit (jobs ~6 GB of 26); the
cores were oversubscribed until the system's own watchdog starved. Lost: E183 at SCF
iteration 1 (nothing), E173 ~3 h into seed 0, E173c ~1.5 h, one suite run (its ".FFF.F"
tail is thrash, not a regression — to be re-run in a free slot). Intact: every record,
the DFT store (9 files), ece_v4_pureref, every finished verdict. No result recorded above
depends on a job that died. Rule now in OVERNIGHT.md and in memory: pw.x ≤ 6 ranks,
ONE search beside it, the suite only in an empty slot, everything nice -n 10, nothing
launched above load 14. Also corrected: my hand-written trace times had drifted up to four
hours from the clock; traces are stamped from date from here on (2026-09-18 09:16).
Colab CPU smoke (2026-09-18 09:44) — measured, and the answer is no for the searches as written.
Operator: "use the colab cli". One CPU session (colab new): 2 cores Xeon 2.20 GHz, 12 GB,
Python 3.13. Payload 100 MB (code + the six files a search reads); a single 100 MB upload
broke the CLI's tunnel in under a second, 12 parts of 8 MB uploaded cleanly and
reassembled to the same sha256. icet, ase, scipy installed on 3.13 without trouble;
FORAGER_NO_CALC=1 (the searches never call MACE). E171b's protocol, 4 rounds × 4 × 1 seed =
16 verifications: not finished after 25 minutes with both cores saturated (load 2.8);
the Mac does 16 verifications in about 3. The hot loop is a sparse mat-vec over 25.1 M
synapses, memory-bandwidth-bound; the VM is ≥ 8× slower at it. One 800-verification seed
would take ≥ 20 h — beyond the session limit and far beyond what the CLI's handle survives
(1 h token, keep-alive). Session stopped; nothing left running on the account.
What Colab is good for here: GPUs. The same mat-vec as a torch sparse CSR product on a
T4/L4 should beat the M4, not trail it — that needs an optional device path in the
propagation (_states, forward) with an equivalence test against scipy (scores equal
to 1e-6 on CPU torch) before any GPU run. Also the eCE fits and, later, MACE fine-tuning.
Not built tonight.
EXPERIMENTS.md · lines 11811–11836
E183 result (13:02) — prediction 1 CONFIRMED, and the cause is not what E181 said.
Hf₁₄Zr₁₃Ti₈Nb₁₁Ta₈ at the standard, 22 iterations: E_form(DFT) = +89.7 meV/atom vs v4 −35.1
— wrong sign, 125 off. Two hcp-former-rich cells (E172: +100 vs +12; E183: +90 vs −35)
make the edge failure regional. But it is not extrapolation: RHEA holds 554 training
cells with Hf+Zr+Ti > 0.5, and their corrected labels average −47 meV/atom. The eCE
reproduces its labels; the labels disagree with DFT by ~140 in that region. E181's
"the model's own extrapolation failure" is withdrawn.
The mechanism, measured (references implied by e_lattice − e_formation, least squares):
el ref used lowest raw bcc pure ref − raw how the ref was set
Mo −10.9154 −10.8503 −65 meV corrected pure cell (correct: rattle removed)
Ta −11.7940 −11.7490 −45 corrected pure cell
Hf −12.3721 −12.5219 **+150** fitted intercept
Ti −7.6645 −7.6837 +19 fitted intercept
Zr −8.1505 −7.9579 **−193** fitted intercept
A corrected reference sits below the raw rattled minimum by the rattle energy (Mo, Ta:
45–65 meV). Hf's fitted intercept sits 150 meV above it — 200–300 meV above where a
corrected bcc-Hf cell would be — so every Hf-rich formation energy is pushed favourable by
~0.2–0.3 × x_Hf eV; Zr's is 190 below, Ti's near neutral. The --pure-references pass
fixed the six refractories from corrected pure cells and fitted the three hcp-formers;
that choice is the defect. Branch → E186: corrected bcc pure references for Hf/Ti/Zr
from RHEA's own bcc pure cells (16 / 6 / 3 exist), all nine fixed, refit → ece_v5.
Prediction: v5 re-predicts E172 and E183 within 30 meV of +100 / +90, the refractory four
stay within 0.8–1.3, and the holdout slope stays ≥ 0.8.
Related entries
E172 — the first rung-4 verdict: DFT vs the cheap rung on one 54-atom cell
E164 — Mixing VASP and Quantum ESPRESSO data: the prior art, and the mechanism it fixes
E173 — serotonin, taught: E171b's protocol with the confirm rung on
E173c — result (2026-09-19 20:23) — the serotonin half of N6, legacy pair