OPTIMIZER AUTOPSY forks training at a detected pre-spike step and intervenes directly on Adam's state, showing where in (w, m, v) the damage lives, when it can be surgically repaired to recover the pre-spike trajectory, and when it provably cannot — turning a folklore mitigation into a predictable, measurable operation.
The instrument (C1 — deterministic replay, snapshot, and the fork Δ==0 gate) is built and
CUDA-verified on a free Kaggle T4, with run evidence committed under results/. The
project is now restarting on a dedicated AMD MI300X budget (Path B: port everything, re-earn only
the ROCm determinism guarantee via a cheap go/no-go smoke test first). Venue shifts to TMLR /
NeurIPS 2027 — NeurIPS 2026 has passed. This page reflects PLAN V6 throughout (the earlier
16-week, team-scaled, A100-based V3 plan has been superseded on hardware, team, budget, and timeline).
Full plan: PLAN_V6.md.
Identity: causal localization + repair — not predict-and-repair. The abstract's one-liner: "We fork training at the moment of failure and intervene directly, showing where in the optimizer the damage lives, when it can be surgically repaired to recover the pre-spike trajectory, and when it provably cannot." Everything else supports that sentence. (SNR-SURGEON survives as the localizer inside C2, demoted from oracle to instrument.)
A fork-and-intervene protocol: snapshot (w, m, v) at a detected pre-spike step, run matched branches under different surgical interventions plus controls — same data order, same seed — and measure divergence in final converged loss. Nobody does controlled interventional training-run science at scale; the closest work stitches checkpoints observationally (LLM360 K2). A reusable scientific instrument. Status: the harness, snapshot, and fork Δ==0 gate are done and verified on a Kaggle T4 (evidence in results/); the guarantee is being re-earned on AMD next.
Use C1 to causally attribute divergence across w, m, v, and test whether the damage is low-rank / subspace-structured or delocalized. Directly resolves the puzzle SPAM's moment-reset ablation raised. Directional SNR from microbatch disagreement is repurposed as the localizer/trigger — "which subspace is poisoned" — not a future-predicting oracle.
Curvature/SNR-guided repair that rescales the identified poisoned subspace of v toward its pre-spike EMA while preserving descent information in the complement. Benchmarked against SPAM, ZClip, AdaGC, global reset, and the honest cheap baseline (spike-skip + clip) on recovery of final loss — not merely "survived divergence."
Spike onset is taken (arXiv 2506.04805). We prove a localizability criterion for recovery: when rank-k repair provably matches global reset, when bulk poison mass makes reset provably necessary, and the computable crossover between them. Theory that predicts its own failure regime.
"Interventional attribution" in the sandbox; "causal" reserved for where it's earned. Natural spikes included (data-corruption class + LLM360 K2's released spike checkpoints), not just induced ones. Commutator correction measured, not assumed away. 3 seeds at 410M stated plainly.
If selective repair never beats global reset → pivot to "the poison is delocalized: a causal explanation of why global reset wins" — still a real result.
If final-loss recovery ≈ cheap spike-skip → the method is dead; ship C1+C2 as a science/benchmark paper to NeurIPS D&B.
Knowing the fallback is what makes this fundable.
| # | Resource | Why |
|---|---|---|
| 1 | 3Blue1Brown — Essence of Linear Algebra (eigenvector chapters, esp. Ch. 14) | Eigen-intuition everything else builds on |
| 2 | 3Blue1Brown — Neural Networks series (backprop chapters) | Gradient mechanics |
| 3 | Karpathy — Zero to Hero: "Let's build GPT" + "Let's reproduce GPT-2 (124M)" | The 124M video is your Week-2 infra tutorial. Watch twice. |
| 4 | Jeremy Cohen — Edge of Stability talks (YouTube: "Cohen Edge of Stability" — Simons Institute / ML theory seminars) | The stability regime the theory borrows |
| 5 | Any good "Adam optimizer explained / bias correction derivation" lecture + YouTube: "loss spikes LLM training" | You will be operating on m and v with a scalpel — know them cold |
autograd.grad(create_graph=True) for HVPs · PyHessian · nanoGPT (fork base) · ZClip official repo (baseline) · SPAM official implementation (linked from arXiv 2501.06842) · Weights & Biases — every run logged, non-negotiableOne shared Zotero library. Every paper gets a 5-line note: claim / method / what we borrow / how we differ / cite-where. Week 1: fresh arXiv sweep for 2025–26 spike papers. The onset-theory paper (2506.04805) must be cited prominently and the recovery claim staked sharply against it — same regime, different question.
| Paper | arXiv | Why it matters |
|---|---|---|
| Adaptive Preconditioners Trigger Loss Spikes in Adam (2025) | 2506.04805 | The onset theory — v decouples from squared gradients, preconditioned sharpness crosses threshold, five-stage spike anatomy. Our theory starts where this stops: recovery, not onset. Cite prominently. |
| Huang et al. — SPAM: Spike-Aware Adam with Momentum Reset | 2501.06842 | The named puzzle C2 resolves: their ablation shows moment reset matters but not where or why. Primary baseline + primary foil. |
| Kumar et al. — ZClip: Adaptive Spike Mitigation for LLM Pre-Training | 2504.02507 | Baseline — z-score EMA anomaly detection on gradient norms; reimplement from their repo |
| AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping | 2502.11034 | Baseline — per-parameter adaptive clipping; also the catalog of natural spike causes (data, hardware, precision, hyperparameters) our spike-sourcing section must answer |
| Molybog et al. — A Theory on Adam Instability in Large-Scale ML | 2304.09871 | The original "poisoned moment buffer" story — foundation of the repair framing |
| Wortsman et al. — Small-scale Proxies for Large-scale Transformer Training Instabilities | 2309.14322 | Spike-induction methodology; we inherit its legitimacy but must additionally validate mechanism transfer across our own two scales |
| Paper | Where | Why |
|---|---|---|
| LLM360 — K2: Building a 65B 360-Open LLM (spike sections + released spike checkpoints) | 2501.07124 | The observational version of what we do interventionally; their public spike checkpoints are our natural-spike test set |
| Ma et al. — Understanding Silent Data Corruption in LLM Training | 2502.12340 | The SDC case: a spike near the end of fine-tuning → zero test accuracy. Proof that "reaction is enough" fails somewhere — the regime C1 must find |
| Cohen et al. — Adaptive Gradient Methods at the Edge of Stability | 2207.14484 | Preconditioned sharpness ≈ 38/η — the working eigenbasis for the localizer |
| Cohen et al. — GD Typically Occurs at the Edge of Stability | 2103.00065 | EoS foundations |
| McCandlish et al. — An Empirical Model of Large-Batch Training | 1812.06162 | Gradient noise scale / SNR formalism the localizer generalizes directionally |
| Damian, Nichani, Lee — Self-Stabilization: The Implicit Bias of GD at EoS | 2209.15594 | Implicit projected-GD machinery reused in the repair analysis |
| Arora, Li, Panigrahi — Understanding GD on the Edge of Stability | 2205.09745 | Quadratic-regime proof techniques |
| Chowdhery et al. — PaLM (loss-spike sections) | 2204.02311 | The restart-and-skip folklore recipe we turn into a measured operation |
optimizer-autopsy/
harness/ # fork/snapshot/replay — the product
localizer/ · repair/ · theory/
experiments/llm/ (nanoGPT fork) · experiments/proxy/
baselines/ (SPAM · ZClip · AdaGC · clip · skip)
analysis/ · spikes/ (recipes + K2 checkpoints)
Deterministic replay is a first-class requirement: seeded loaders, RNG state in every snapshot, bitwise-identical trunk verification in CI. A 10-minute smoke test forks a tiny run and asserts branch divergence is zero under no-op-with-no-spike. Reproducibility isn't a virtue here — it's the instrument's calibration.
LLM track: OpenWebText or FineWeb 10BT sample (HuggingFaceFW/fineweb) on GPT-2 124M and Pythia-style 410M. You need spike windows, not convergence: 2–5B tokens per 124M run.
Induced spikes: high LR, tiny Adam ε (1e-12), bf16→fp16 on sensitive ops, weakened QK-layernorm.
Natural spikes (mandatory): data-injected corrupted-batch spikes — the cheapest natural class (AdaGC's catalog: data quality, hardware faults, precision, hyperparameters) — plus LLM360's released K2 spike/normal checkpoint pairs as a real-world validation set. A repair that only handles ε-perturbation spikes at 124M is a toy.
The earlier "500–600 GPU-hours" figure conflated two different resources: build-effort (person-hours) and GPU-compute. A bottom-up recount — grounded in a measured 0.0619 s/step at proxy scale (11M params, batch 16, Tesla P100; committed under research/kaggle/step_timer_results.md) and a full code audit of which components are actually GPU-bound versus engineering-time — puts the real AMD GPU ask at ~15–40 GPU-hours. That is ~5–10h to re-earn bit-exact determinism on AMD/ROCm (the one line nothing has tested — every measurement so far ran on NVIDIA/CUDA free-tier Kaggle) plus ~3–30h of proxy-scale GPU-bound science (calibration + attribution battery). The ~500 remaining hours are build-effort, which drives the ~24-week wall-clock, not the GPU grant. Path B: port every hardware-independent piece (data, model, spike recipes, fork/branch design, and the already-verified determinism/snapshot/fork spine) unchanged; re-earn only the ROCm determinism guarantee. A cheap go/no-go smoke test runs first — one forked pair, one step, compare m/v/w bit-for-bit over the exact ops this pipeline uses (HVP double-backward + Adam moment update) — because ROCm has no exact CUBLAS_WORKSPACE_CONFIG analog and bitwise Δ==0 on MI300X is not assumed.
Bottom-up recount. Build-effort = person-hours (drives wall-clock, not the GPU grant). GPU-h (audited) = real proxy-scale compute at the measured 0.0619 s/step, low–high; the assumption behind each range is stated. Rows marked ENG are engineering-bound (their GPU cost is only dev-validation runs); GPU rows are genuinely compute-bound.
| Line item | Build-effort (h) | Bound by | GPU-h (audited) | Basis |
|---|---|---|---|---|
| AMD determinism re-verification | 30 | ENG + AMD | 5–10 | The one unverified line. Bit-exact Δ==0 has only ever run on NVIDIA/CUDA; re-earning it on ROCm is the go/no-go. Small compute, high risk. |
| Spike induction + detection | 30 | ENG | 0.1–1.4 | induce + detector built & run (1/4 recipes pass DoD); remainder is coding + threshold sweeps over cached runs, not retraining. |
| Cheap-fix kill-test | 15 | ENG (built) | 0.1–0.5 | Battery + Gate-B verdict built, not yet run: 4 branches × (2–4 recipes) × 3–5 seeds × 200 steps. |
| Localizer (C2) | 70 | ENG | 0.1–2 | Pure TODO stubs today; HVP/eigensolve on the 11M proxy is cheap — the 70h is coding, not GPU. |
| Repair operator (C3) | 60 | ENG | 0.2–1.7 | 9-line stub; repair is a cheap tensor edit validated by 50–200 short forks vs the clean counterfactual. |
| Baselines | 45 | ENG | 0.2–1.3 | skip/clip/reset real; SPAM/ZClip/AdaGC are 3-line stubs — reproducing published methods is human time. |
| Short-fork calibration | 55 | GPU | 2–14 | Genuinely GPU-bound: ~20 full-length proxy forks (5000 steps) to fix the shortest fork length that preserves branch-ordering. |
| Full attribution battery (proxy) | 120 | GPU | 0.4–4.8 | recipes × 7 branches × seeds × fork-len × 0.0619s ÷ 3600: (2×7×3×500)→0.4h … (4×7×5×2000)→4.8h. A ceiling. |
| Commutator-error measurement | 15 | Analysis | 0.1–1 | Measured once at 1–4 confirmed spike sites (a few HVPs each). |
| Contingency | 60 | Buffer | — | Effort buffer, not GPU compute. |
| Total | ~500 | — | ~15–40 | The ~500 is build-effort (PLAN V6 §5 flagged the GPU-vs-dev-hours conflation). The audited AMD GPU ask is ~15–40 GPU-h. |
The larger-scale (124M) robustness battery measures ~71–302 GPU-hours per battery (measured 124M step-time 1.08 s/step at batch 8 on P100, grad_accum=40; see step_timer_results.md §2). It is designed to run on free-tier Kaggle over multiple weeks per BUILD_PLAN Task 22's own weekly-cap survival design — not on the AMD grant, and it is not part of what is requested from Exea Labs.
Written cut order (decided now, not mid-crunch): (1) shrink to the recipes that work; (2) fewer seeds, flagged as a limitation; (3) drop the two adaptive-clipping baselines; (4) measure the commutator error once, not per site. Never cut: the exact-zero replay proof or the held-out validation methodology. Snapshots: proxy (w,m,v) fp32 ≈ 12–36 MB, 124M bf16 ≈ 0.75 GB → rotating latest.safetensors on HuggingFace. Underlying data + full audit: research/kaggle/step_timer_results.md; plan detail: PLAN_V6.md.
Scheduling uses a conservative single-contributor-plus-mentor assumption. The team-size figure is a PLAN V6 open item: different numbers have circulated (five with a mentor; eight per outreach framing) against a commit history showing one author. It needs to resolve to one accurate figure before it appears in any funding or compute request — cheap to fix now, costly to credibility if a partner checks it later. If a verified team with assigned roles is confirmed, the localizer and repair builds can run in parallel and the wall-clock timeline compresses; the GPU-hour budget itself does not change.
| Work area | State | Scope |
|---|---|---|
| Instrument (C1) — harness, snapshot, fork Δ==0 gate | ✅ done, CUDA-verified | Deterministic replay, name-keyed (w,m,v,RNG) snapshot/restore, fork-and-compare. Re-earn on AMD via the go/no-go gate. |
| Spikes + kill-test | 🟡 partial | 1 of 4 induced-spike recipes solid; cheap-fix battery built, not yet run to a verdict. |
| Localizer (C2) + repair (C3) + baselines | ⬜ not started | Directional SNR + preconditioned-Hessian curvature → poison score; rank-limited projection repair; SPAM/ZClip/AdaGC/skip+clip/reset baselines. The genuinely novel, highest-risk work. |
| Theory | ⬜ not started | AR(2) recovery criterion + ψ_k crossover; commutator-error measurement; the per-spike falsifiable prediction. Mentor: proof checking, venue strategy, endorsement. |
Compute (open items, from PLAN V6 §5): the audited AMD GPU ask is ~15–40 GPU-hours (see §04) — the earlier 500–600h figure was build-effort, not GPU-compute. Still to confirm with Exea Labs: the realistic single-project AMD allocation, plus GPU parallelism + any grant expiry, since those decide whether the (small) GPU-hours or personal dev-hours is the binding constraint. Nothing is locked until the Exea inquiry resolves.
The per-phase hour tags below are build-effort (person-hours), not GPU-hours — the audited GPU-compute ask is the separate ~15–40 GPU-h in §04. Wall-clock is the single-contributor-plus-mentor pacing (~24 weeks); it compresses if a verified team or parallel GPUs materialize. The week numbers are relative to the AMD grant landing, not fixed calendar dates. Venue: TMLR (no deadline, judges claim-support) and/or a NeurIPS 2027 cycle — NeurIPS 2026 has passed, and 2027's deadline is a historically-grounded prediction (2024 May 22 → 2025 May 16 → 2026 May 6), not a locked date.
No reallocation of a fixed compute budget engineers a high acceptance rate — review is noisy even for strong papers, and the top tier needs a different category of contribution than a well-executed application of existing tools to a scoped question. The honest ceiling here is a genuinely strong, technically sound single contribution: exactly what TMLR is built to reward, and a reasonable NeurIPS D&B or workshop candidate. Not a near-certain top-tier acceptance, and this plan does not claim to be.
| Artifact | When | Venue | Fit |
|---|---|---|---|
| Preprint (C1 + C2 + theory sketch) | After the attribution battery | arXiv | Direct |
| Full paper | ~Weeks 21–24 | TMLR (rolling; judges claim-support, no deadline) | Strong structural fit |
| Full paper (alt) | NeurIPS 2027 cycle (predicted ~May 2027) | NeurIPS 2027 main / D&B | Reasonable |
| If the method dies (kill-test) | Decided in Phase 2 | NeurIPS D&B — the harness + causal benchmark | Reasonable |
| NeurIPS 2026 main | Out of reach — the May 6, 2026 deadline has passed. | — | |
Three tiers (Floor / Core / Stretch) with a written cut order, so a smaller-or-larger grant needs no renegotiation. The kill-test and go/no-go run first, so months aren't spent on a dead or unprovable contribution.
Training counterfactuals as an instrument: no direct prior work found. C2 resolves a named, published puzzle (SPAM's ablation). C3 is "beat reset" — incremental, never the lead. Write the paper in exactly this order.
Repair works → C1+C2+C3; cheap fix wins → C1+C2 benchmark; poison delocalized → "why reset wins." Every branch is a real paper. "Causal" only where the random-subspace control earned it; pre-registered thresholds remove the pull toward the exciting read.
Open logistics (PLAN V6 §5, §11): resolve the team-size figure, confirm the realistic Exea allocation, and confirm GPU parallelism + expiry before locking a tier. Author order, mentor affiliation, and endorsement follow from the team resolution.
Every item below was raised adversarially against this plan. They stay on this page so nobody discovers them in a rebuttal.
The entire evidentiary standard is exact-zero replay, and the CUDA trick that earned it (CUBLAS_WORKSPACE_CONFIG set before torch imports) has no exact ROCm analog; rocBLAS/hipBLASLt determinism is handled differently and deterministic-kernel coverage is less complete than CUDA. If Δ==0 is unachievable on MI300X, the proof standard collapses on that hardware. Mitigation: the ~6h go/no-go smoke test runs first, on the exact ops this pipeline uses; if it fails, Phase 1 pivots to a tolerance-based statistical framework (paired seeds, many-run averaging) rather than discovering the wall weeks in. A useful narrowing: the eigensolve appears to run its iterative solve on the host and farm only matvecs to the GPU — so the burden may be one bit-reproducible HVP, not a whole multi-step solve. Confirm that.
The AMD restart is contingent on an Exea Labs grant that is requested, not approved. The audited ask is modest — ~15–40 GPU-hours (the old "500–600 GPU-hours" was build-effort mislabeled as compute) — so the binding constraint is more likely personal dev-time than GPU-hours, but "24 weeks" and any hour figure are only the same constraint if there's one GPU with no expiry. Mitigation: three scope tiers (Floor/Core/Stretch) so a smaller-or-larger grant needs no rewrite; PLAN V6 §5 lists the exact questions to resolve (realistic allocation, parallelism, expiry) before locking a tier. Abandoning a working CUDA stack for an unconfirmed one is the real cost if this falls through.
If the kill-test reveals that one-step-late spike-skip recovers held-out loss as well as any surgery, the repair contribution collapses — a live possibility, since the field skips-and-reinjects routinely. Mitigation is baked in: the cheap-fix battery runs early (Phase 2), against a pre-registered threshold, so the answer is known before the expensive machinery is built. If reaction wins → C1+C2 ship as a science/benchmark paper; the methodology de-risks everything either way.
Per-direction decoupling assumes the preconditioner commutes with the Hessian eigenbasis; it doesn't, and the onset paper (2506.04805) already lives in this regime. Mitigation: frame the decoupled result as the leading-order characterization, measure the commutator's effect on ψ* empirically, and claim a leading-order theory with a measured correction. A measured correction is more credible than a clean-but-false bound. PLAN V6 §8 pre-registers what a convincing ψ_k signal is before it decides anything.
Fork-and-intervene with matched seeds and a random-subspace control is genuinely interventional — far stronger than observational checkpoint stitching — but it's causal within the induced-spike sandbox; transfer to natural spikes is inference. Mitigation: say "interventional attribution," reserve "causal" for the sandbox, and let the K2 + corrupted-batch results carry the transfer argument.
C1 (training counterfactuals as an instrument): freshest thing here, no direct prior work found, and it's already built and CUDA-verified. C2: resolves a named, published puzzle — reviewers value that above generic gaps. C3: "beat reset," incremental, never the lead. Theory: differentiated (recovery vs onset) but shares a regime with 2506.04805 — cite it prominently, stake the recovery claim sharply. Net: a defensible novelty profile whose strength is the methodology and the puzzle-resolution, not the optimizer tweak.