From de50b102a7863c77c68f0f57895bb73c6614f420 Mon Sep 17 00:00:00 2001 From: odunbar Date: Wed, 8 Jul 2026 17:15:19 -0700 Subject: [PATCH 1/2] stop global memory --- .claude/settings.json | 1 + 1 file changed, 1 insertion(+) diff --git a/.claude/settings.json b/.claude/settings.json index 4990e57e8..153fe4042 100644 --- a/.claude/settings.json +++ b/.claude/settings.json @@ -1,5 +1,6 @@ { "$schema": "https://json.schemastore.org/claude-code-settings.json", + "autoMemoryEnabled": false, "permissions": { "allow": [ "Bash(julia --project *)", From 93763bfedda0decde9e848f5dcee4d940077f6ad Mon Sep 17 00:00:00 2001 From: odunbar Date: Thu, 9 Jul 2026 12:57:51 -0700 Subject: [PATCH 2/2] math-auditor skill --- .claude/skills/math-auditor/SKILL.md | 229 ++++++++++++++++++ .../references/unit-agent-prompt.md | 87 +++++++ 2 files changed, 316 insertions(+) create mode 100644 .claude/skills/math-auditor/SKILL.md create mode 100644 .claude/skills/math-auditor/references/unit-agent-prompt.md diff --git a/.claude/skills/math-auditor/SKILL.md b/.claude/skills/math-auditor/SKILL.md new file mode 100644 index 000000000..be3b9effb --- /dev/null +++ b/.claude/skills/math-auditor/SKILL.md @@ -0,0 +1,229 @@ +--- +name: math-auditor +description: > + Run an adversarial mathematical-accuracy review of a Julia package's src/ and + test/ directories, producing a dated markdown report plus concise, self-contained + fix-prompt markdowns suitable for handing to a smaller model (e.g. Sonnet) in a + later session. Use this skill whenever the user asks for an adversarial review, + a math audit, a full code review focused on mathematical/numerical correctness, + a check of algorithmic consistency with the literature, or says things like + "review the math in src/", "is the algebra right?", "audit the update equations", + "check the statistics/linear algebra for bugs", or "construct a code review as + markdown". Trigger even when the user does not say "audit" — any request for a + correctness-focused sweep of a scientific Julia codebase qualifies. +--- + +# Math Audit + +Adversarial review of a scientific Julia package for **mathematical accuracy and +consistency** — not software architecture (flag architecture only when it causes +mathematical wrongness, e.g. mutation aliasing, accidental type demotion, or +inconsistent conventions between modules). + +The output is written for the package's own developers: findings must cite exact +`file:line`, state the correct mathematics, and give a concrete failure scenario. +A finding that can't survive an attempt at refutation doesn't ship. + +## What "adversarial" means here + +Each reviewer's job is to *break* the code, not describe it. Concretely, hunt for: + +- **Wrong equations**: emulator predictive mean/covariance formulas, kernel and + feature-map definitions, MCMC proposal/acceptance ratios, likelihood and prior + terms that differ from the cited papers or from the docstring's own LaTeX. + Derive the correct expression independently and diff it against the code. +- **Convention drift**: rows-vs-columns for ensemble members/input points, + `N-1` vs `N` normalization, factor-of-2 / sign errors, Cholesky `L` vs `U`, + covariance vs precision, whitened vs unwhitened space, whether decorrelation + is applied in input- or output-space — especially *inconsistencies between + modules that must agree* (e.g. `GaussianProcess.jl` vs `ScalarRandomFeature.jl` + vs `VectorRandomFeature.jl` on the same emulator interface). +- **Statistical validity**: is observational/model noise sampled and scaled + correctly? Are means/covariances computed over the right dimension (samples + vs output components)? Do scalar and vector emulator variants agree in + expectation? Do the SVD/PCA-based decorrelation and elementwise scaling + transforms in `Utilities/` compose correctly with the emulator and sampler + that consume them? +- **Numerical soundness**: unguarded `inv`/`\` on possibly-singular matrices, + loss of symmetry/PSD-ness, subtraction-based variance formulas, missing + regularization/jitter, `sqrt` of negative-by-roundoff eigenvalues. +- **Edge cases the math must survive**: single training point, single output + dimension (scalar vs matrix degeneracy), zero variance, degenerate/rank- + deficient covariance after decorrelation, empty or singleton minibatches, + MCMC chains that reject every proposal. +- **Test-math consistency**: do the tests actually pin the mathematics + (analytic solutions, invariants, convergence rates), or just check shapes and + "it runs"? A wrong equation whose test only checks `size()` is a *double* + finding: the bug and the missing test. + +## Workflow + +### 1. Partition + +First establish what is actually **live**: grep the top-level module (and the +files it includes, recursively) for `include(...)`, and compare against +`git ls-files src/`. Files never included in the build, or untracked, are dead +code — assign them skim-only status and cap their findings at minor/hygiene +(they can't produce wrong results today, but note them in the report so future +reviewers and greps don't confuse dead sampler math with live code). + +List the live `src/*.jl` and `test/**/*.jl` with line counts. Group into 4–8 +review units of roughly comparable size, pairing each source module with the +tests that exercise it. Group modules that *share mathematical conventions* +together so the reviewer can catch cross-module inconsistencies — in this +package that typically means: the emulator variants (`Emulator.jl` plus +`MachineLearningTools/GaussianProcess.jl`, `RandomFeature.jl`, +`ScalarRandomFeature.jl`, `VectorRandomFeature.jl`) in one or two units since +they must agree on the same predict/covariance interface; the MCMC samplers in +another; and the `Utilities/` preprocessing transforms (decorrelation, +elementwise scaling, canonical correlation, likelihood-informed subspace) in +another, since they feed both the emulator and the sampler and a convention +mismatch there propagates silently downstream. + +Reserve the smallest unit as a **cross-cutting unit**. Alongside its own files +(module wiring, display code, package-wide greps for `dims=`, `corrected=`, +jitter constants, symmetrization idioms), give it an explicit mandate to +independently re-derive the *single highest-risk formula* in each other unit's +territory — MCMC acceptance ratios, structure-matrix congruence transforms, +predictive-covariance noise semantics. This overlap is deliberate: in one +audit the cross-cutting unit independently found and numerically verified the +same critical pCN acceptance-ratio bug as the MCMC unit, on a *different* toy +problem — two matching numerical confirmations settle a critical finding in a +way a single one never can. Tell that agent the overlap is intentional so it +doesn't defer to the "owning" unit. + +### 2. Fan out reviewers (parallel agents) + +Spawn one agent per unit, in a single message so they run concurrently. Build +each prompt from `references/unit-agent-prompt.md`: fill the placeholders +(unit file list, domain context, unit-specific extra hunts, scratch directory) +and paste in the full "What adversarial means here" list above — agents don't +see this skill. The template carries the rules that have earned their keep in +past audits (the numerical-verification tiering, the warn/throw grep before +any "silent" claim, severity-rule reporting) and fixes the required +three-section output: the findings JSON, a CONVENTIONS section, and a +TEST-COVERAGE section. The CONVENTIONS lines are the raw material for the +step-3 matrix — without them you must re-read the code yourself to build it. + +As each agent's report arrives, save its raw findings verbatim to a scratchpad +file (one per unit). A full audit plus verification is long enough that +context summarization mid-run can silently lose findings; the scratchpad files +are the durable record the report is assembled from. + +### 3. Verify + +For each critical/major finding, attempt refutation before it enters the +report: re-read the cited lines yourself, re-derive the math, and check whether +a test or an upstream transformation already accounts for it (common false +positives: a transpose hidden in a helper, normalization done at construction +time, a convention documented elsewhere, a `@warn` at the constructor that the +reviewer never read — re-run the warn/throw grep yourself for any "silent" +claim). Prioritise findings tagged `verified: inspection`; numerically verified +ones usually need only a sanity re-read. Spawn skeptic agents for findings you +can't settle from the main context. Demote or drop findings that don't survive; +mark surviving ones **CONFIRMED** vs **PLAUSIBLE** (couldn't fully verify). + +Then build a small **conventions matrix** from the agents' CONVENTIONS +sections before writing anything: rows = modules, columns = the conventions +they reported (covariance normalization N vs N−1, whitened vs unwhitened +space, whether predictive covariance includes observational noise, +symmetrization idiom, jitter constants, RNG threading, rows-vs-columns). Any +mismatched cell between modules that must agree is a finding candidate in +itself — in practice the worst bugs are a scale factor or transform applied in +one emulator variant or preprocessing step but not its sibling, and they only +become visible side by side. + +### 4. Write the report + +Create `full-code-review//` (date from `date +%F`, never from +memory). Before writing, copy every numerical-verification scratch script into +`full-code-review//verification//` and cite those repo-relative +paths in the report — scratchpad paths die with the session, and a numerical +finding whose evidence has evaporated is just an inspection finding with extra +steps. + +Write `review.md`: + +```markdown +# Adversarial Mathematical Review — () +## Scope and method +## Summary table +## Critical findings +## Major findings +## Minor findings & hygiene +## Cross-module consistency notes +## Test-coverage gaps +## What was checked and found sound +``` + +In **Test-coverage gaps**, lead with the 2–3 *keystone tests* whose absence +masked the most findings, ranked by findings-caught-per-test. In practice a +handful of missing test archetypes hide many bugs at once — e.g. an analytic +linear-Gaussian posterior test asserting the MCMC mean **and variance** would +have caught two critical sampler bugs in one audit, and a single test of the +default noise-flag semantics would have caught three emulator findings. This +list is the highest-leverage part of the report for maintainers. + +Findings get stable IDs used everywhere, including fix prompts: `C1…` critical, +`M1…` major, `m1…` minor, `h1…` hygiene; grouped-minor fix prompts get `G1…`. + +### 5. Write fix prompts + +For each finding with an actionable fix (usually critical + major, plus grouped +minors), write `full-code-review//fix-prompts/-.md`. These are +consumed by a *smaller model in a fresh session with no context*, so each must +be self-contained: + +```markdown +# Fix : +**File**: `src/Foo.jl`, function `bar!`, around line NNN. +**Problem**: <2–4 sentences: what the code does vs what the math requires. +Include the incorrect snippet verbatim.> +**Required change**: +**Do not**: +**Verify**: +``` + +Keep each under ~40 lines. One finding per file; group only truly mechanical +repeats (e.g. the same typo pattern in five docstrings) into one prompt. +Quote enough of the offending snippet that the fixer can locate it by function +name + snippet — line numbers drift between the audit and the fix session, so +present them as hints, not anchors. +Also write `fix-prompts/README.md` listing prompts in recommended application +order (independent fixes first, same-file prompts sequenced, conflicting ones +flagged) and naming which fixes *intentionally change numerical results* — so +the fixer checks a failing loose regression test against the analytic reference +in the prompt before "fixing" the test. + +### 6. Report back + +Final message: lead with the headline (how many confirmed critical/major +findings and the single worst one), then the report path, then a compact +summary table. Do not paste the whole report into the chat. + +## Calibration + +- Severity: **critical** = produces mathematically wrong results in mainstream + use; **major** = wrong in common configurations or silently degrades + statistical properties; **minor** = wrong in edge cases, misleading docs + math, dead/misnamed math; **hygiene** = style-level (only if math-adjacent). +- A numerically demonstrated wrong stationary distribution, wrong acceptance + ratio, or wrong variance scaling in *exported* functionality is **critical** + even if the feature looks niche — samplers and estimators exist to produce + distributions, and a wrong distribution is the worst failure they have. + (Agents report which calibration rule they applied, so when two units rate + the same bug differently you arbitrate on the rule, not the adjective.) +- A docstring–code mismatch is a real finding even when the code is right — + users implement against docstrings. +- A statistically unjustified combination that the package explicitly warns + about (e.g. a constructor `@warn "... experimental ..."`) caps at **minor** + unless the warning itself is wrong — the trap isn't silent. +- Don't pad. If a unit is sound, the report says so in one paragraph; an audit + that cries wolf gets ignored next time. + +## Improving this skill + +After delivering the report, offer: "Would you like to improve the +**math-auditor** skill itself using skill-creator? You can share suggestions, or +I can analyse this run — finding quality, false-positive rate, fix-prompt +usability — to refine the skill for next time." diff --git a/.claude/skills/math-auditor/references/unit-agent-prompt.md b/.claude/skills/math-auditor/references/unit-agent-prompt.md new file mode 100644 index 000000000..9e735461c --- /dev/null +++ b/.claude/skills/math-auditor/references/unit-agent-prompt.md @@ -0,0 +1,87 @@ +# Unit-agent prompt template + +Fill the `{PLACEHOLDERS}` and send one such prompt per review unit, all in a +single message so the agents run concurrently. `{HUNTING_LIST}` is the full +"What adversarial means here" bullet list from SKILL.md, pasted verbatim — +agents cannot see the skill. `{EXTRA_HUNTS}` is where unit-specific attack +targets go (derive them from the unit's domain: e.g. for an MCMC unit, spell +out the correct pCN/Barker/MALA acceptance-ratio forms; for preprocessing +units, the congruence-transform and round-trip invariants; for the +cross-cutting unit, the deliberate-overlap mandate and the package-wide greps). + +--- + +You are an adversarial mathematical reviewer for the Julia package {PACKAGE} +at {REPO_PATH}. Your job is to BREAK the code mathematically, not describe it. +Focus on mathematical accuracy and consistency, NOT software architecture +(flag architecture only when it causes mathematical wrongness, e.g. mutation +aliasing, accidental type demotion, inconsistent conventions). + +YOUR UNIT (read every line of these): +{FILES — with line counts, and which test files pair with which source files. +Mark any dead/not-in-build files as skim-only.} + +Context: {DOMAIN_CONTEXT — 3-6 sentences: what the package does, what this +unit's math is supposed to compute, the key papers/algorithms it implements, +and how this unit's outputs are consumed by the rest of the pipeline. Name +adjacent files the agent may read for interface contracts but that another +reviewer audits in depth.} + +HUNT FOR: +{HUNTING_LIST} +{EXTRA_HUNTS} + +RULES: +- Prefer few, well-evidenced findings over many speculative ones — but do + report genuine minor inconsistencies. If the module's math is correct, say + so and note the strongest invariants the tests pin. +- Verify cheap claims numerically and tag them `verified: numerical` — + numerically verified findings are worth far more than inspection-only ones. + Tiering, best first: (1) reproduce the behavior in-package + (`julia --project={REPO_PATH}`; a Manifest.toml usually resolves — worth the + precompile wait for behavioral claims like state-sharing, crashes, or + option-ignored bugs); (2) standalone script using only + LinearAlgebra/Statistics/Random/Distributions that reimplements the + questioned formula next to the correct one (right tool for formula and + statistical-scaling claims — e.g. a tiny MH chain on a 1D conjugate Gaussian + exposes a wrong acceptance ratio in seconds); (3) inspection, only when + running code is genuinely expensive. Put scratch scripts in {SCRATCH_DIR} + and cite them per finding. +- Before claiming anything is 'silent' or 'has no warning/guard', grep for + `@warn`, `@error`, and `throw` at the *constructors and call sites* of the + code path, not just the function you are reading — guards often live at + construction time. +- Check code against docstrings/comments AND against the standard form of the + algorithm from the literature. A docstring–code mismatch is a real finding + even when the code is right — users implement against docstrings. +- Severity calibration: critical = mathematically wrong results in mainstream + use — and any numerically demonstrated wrong stationary distribution, + acceptance ratio, or variance scaling in exported functionality is critical + even if the feature looks niche; major = wrong in common configurations or + silently degrades statistical properties; minor = edge cases, misleading + docs math, dead/misnamed math; hygiene = math-adjacent style. A trap the + package explicitly @warns about caps at minor. For each finding, state + WHICH rule you applied. + +OUTPUT (your final message is raw data for synthesis, not prose for a human) — +three sections, all required: + +1. FINDINGS: a JSON-like list, each entry with: `file`, `line`, `severity` + (critical/major/minor/hygiene) plus the calibration rule applied, `claim` + (one sentence), `evidence` (quote the offending snippet vs the correct + math), `failure_scenario` (concrete inputs → wrong output), + `verified` (numerical / inspection, with script path if numerical), + `suggested_fix` (optional, a few lines). + +2. CONVENTIONS: one line per convention this unit assumes — data orientation + (columns = samples?), covariance normalization (N vs N−1), whitened / + encoded-space handling, whether predictive covariance includes + observational noise, symmetrization idiom (`Symmetric` vs `hermitianpart`), + jitter/regularization constants and where they apply, RNG threading. + Explicitly note any DISAGREEMENT between files inside your unit. These + lines feed a cross-module consistency matrix, so state values, not vibes. + +3. TEST-COVERAGE: which mathematical properties the unit's tests pin + (analytic values, invariants, statistical tolerances) vs leave unpinned — + and for each of your findings, whether an existing test could have caught + it (if not, that is a double finding: the bug and the missing test).