From b690eada1c152357675784a47b16571203875566 Mon Sep 17 00:00:00 2001 From: pedrohcgs Date: Wed, 10 Jun 2026 13:56:47 -0400 Subject: [PATCH 1/2] =?UTF-8?q?feat(v2.1):=20currency=20+=20citability=20?= =?UTF-8?q?=E2=80=94=20Fable=205=20refresh,=203=20doc-currency=20bug=20fix?= =?UTF-8?q?es,=20/submission-disclosures,=20CITATION.cff?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Driven by a 48-agent web-verified audit (40 findings, 7.5% noise). Architecture came back clean ("ahead of the field"); every fix here is a FACT fix, not structure. MODEL REFRESH (the false superlative): Fable 5 GA'd 2026-06-09; README + SSoT still said "Opus 4.8 is the newest." SSoT table/marker/date updated per its own protocol; routing rule gains "Where Fable 5 fits — and where it does not" (fleet stays Opus/Sonnet/Haiku: 2x price on the judgment tier + observed 28/28 day-one structured-output protocol failures vs 0 on Opus). check-model-versions.sh hardened: version regex generalized (bare "Fable 5" tracked), Fable tier added, NEW superlative-drift check (the exact class the old gate missed) — which immediately caught + we fixed a statusline comment false-positive by making it version-neutral. DOC-CURRENCY BUGS (all verified against current docs before fixing): - scheduled-routines.md: /schedule flag syntax was FABRICATED -> real natural-language form, /schedule update for cron, 1h min interval, daily cap, committed-repos-only. MCP guardrail was INVERTED: cloud Routines include ALL connectors w/ write access by default -> rewritten to least-privilege ("the risk is a fully-armed connector, not a missing one"). - allowed-tools taught as a sandbox; docs verified (claude-code-guide agent, exact quote): it is a PRE-APPROVAL list — "does not restrict which tools are available". disallowed-tools (the actual restrictor) + paths/when_to_use/arguments now documented in template + guide, with a read-only-skill pattern and a "does not sandbox" warning callout. Security note rewritten. ADDED: - /submission-disclosures (51->52, fully count-wired; guide re-rendered via quarto): AI-use disclosure matched to the journal's VERIFIED-CURRENT policy + CRediT + COI + data-availability. Two independent audits converged on this as the top 2026 norms gap. Not /disclosure-check. - CITATION.cff (citable; Zenodo DOI in backlog), SECURITY.md, CODE_OF_CONDUCT.md. - GitHub topics (10, was zero) + homepage URL set via gh. Backlog: awesome-list PRs, Zenodo DOI steps, hook-touchpoints note (verify names first), persona cost table, non-Claude-coauthor README box. Audience pinned: econ + closely related. Gates: surface-sync (29+2 @ 52 skills), skill-integrity, model-versions (Fable-aware) — green. --- .claude/references/model-versions.md | 12 ++- .claude/references/scheduled-routines.md | 13 +++- .claude/references/v2.0-backlog.md | 9 +++ .claude/rules/model-routing.md | 11 +++ .claude/scripts/statusline.sh | 2 +- .../skills/submission-disclosures/SKILL.md | 76 +++++++++++++++++++ .github/CODE_OF_CONDUCT.md | 7 ++ .github/CONTRIBUTING.md | 2 +- .github/SECURITY.md | 24 ++++++ CHANGELOG.md | 27 +++++++ CITATION.cff | 29 +++++++ CLAUDE.md | 2 +- README.md | 7 +- docs/index.html | 4 +- docs/workflow-guide.html | 51 ++++++++++--- guide/workflow-guide.html | 51 ++++++++++--- guide/workflow-guide.qmd | 22 ++++-- scripts/check-model-versions.sh | 32 +++++++- templates/skill-template.md | 19 ++++- 19 files changed, 354 insertions(+), 46 deletions(-) create mode 100644 .claude/skills/submission-disclosures/SKILL.md create mode 100644 .github/CODE_OF_CONDUCT.md create mode 100644 .github/SECURITY.md create mode 100644 CITATION.cff diff --git a/.claude/references/model-versions.md b/.claude/references/model-versions.md index 0ea5e2e0d..56897fd97 100644 --- a/.claude/references/model-versions.md +++ b/.claude/references/model-versions.md @@ -1,8 +1,8 @@ - + # Current Model Versions (single source of truth) -**Last verified against Anthropic docs:** 2026-05-31 +**Last verified against Anthropic docs:** 2026-06-10 This file is the **one place** that names current Claude model point versions. Everything else in the template should either refer to tiers abstractly ("newest Opus", "the Haiku tier") or point here. `scripts/check-model-versions.sh` flags any **superseded** version that is presented as **current** in the template's user-facing surfaces. @@ -10,7 +10,8 @@ The machine-readable `` marker at the top is parsed by the | Tier | Current version | Model ID | Notes | |------|-----------------|----------|-------| -| Opus (high-judgment) | **Opus 4.8** | `claude-opus-4-8` | newest; API default; GA 2026-05-28; $5/$25 per MTok; 1M context; defaults to `high` effort | +| Fable (Mythos-class; hardest, long-horizon) | **Fable 5** | `claude-fable-5` (alias `fable`; 1M variant `claude-fable-5[1m]`) | most capable model in Claude Code; **opt-in** (`/model fable` or the `best` alias) — NOT the default on any account type; GA 2026-06-09; $10/$50 per MTok; 1M context (128k max output); defaults to `high`; requires Claude Code ≥ 2.1.170; falls back to Opus 4.8 on flagged cyber/bio content | +| Opus (high-judgment) | **Opus 4.8** | `claude-opus-4-8` | current Opus tier; **API/account default**; GA 2026-05-28; $5/$25 per MTok; 1M context; defaults to `high` effort; the content-fallback target for Fable 5 | | Sonnet (workhorse) | **Sonnet 4.6** | `claude-sonnet-4-6` | 1M context | | Haiku (fast / mechanical) | **Haiku 4.5** | `claude-haiku-4-5-20251001` | fast tier; ID is the snapshot-pinned alias | @@ -18,6 +19,8 @@ The machine-readable `` marker at the top is parsed by the **Fast mode:** Opus 4.8 fast mode is $10/$50 per MTok (~2.5× speed). Opus 4.7 fast mode is $30/$150; Opus 4.6 fast mode is deprecated (~2026-06). +**Fable 5 maturity caveat (2026-06-10):** routing guidance lives in [`model-routing.md`](../rules/model-routing.md) — as of launch week, Fable 5 is deliberately **not** routed to the forked-reviewer fleet (2× the Opus price on the judgment tier, and day-one forced-tool-protocol reliability was observed to lag Opus 4.8: 28/28 structured-output subagent failures in one session's workflow vs 0 on Opus). Re-evaluate as point releases land. + ## Prior generations It is fine to mention older versions in **historical** contexts (CHANGELOG entries) or in explicit **"prior generation" / comparison** lines (e.g. "Opus 4.8's `high` does what Opus 4.7's `xhigh` did"). They must **not** be presented as the current / newest / default model. @@ -30,4 +33,5 @@ The checker allows a line to mention an older version when it carries a marker s 1. Update the table **and** the `` marker above, plus the "Last verified" date. 2. Run `./scripts/check-model-versions.sh` and fix every current-state surface it flags. -3. Add a "Changed — model refresh" entry to `CHANGELOG.md`; leave historical CHANGELOG entries intact. +3. **Manually grep for superlatives** — `grep -rniE "newest|most capable" README.md CLAUDE.md guide/ docs/index.html .claude/rules/` — and re-verify each hit. The checker validates *version strings*; a claim like "X is the newest model" is a **semantic** assertion it can only partially catch (it flags `newest`/`most capable` lines that name a non-top tier, but tier-relative phrasings like "the newest Opus" are legitimately allowed). The 2026-06-09 Fable 5 launch made exactly this class of claim false while the gate stayed green. +4. Add a "Changed — model refresh" entry to `CHANGELOG.md`; leave historical CHANGELOG entries intact. diff --git a/.claude/references/scheduled-routines.md b/.claude/references/scheduled-routines.md index ea4edc783..db7ec7dde 100644 --- a/.claude/references/scheduled-routines.md +++ b/.claude/references/scheduled-routines.md @@ -21,18 +21,23 @@ The nightly reproducibility job is the *backstop*. The *immediate* signal is the ## Setting one up +`/schedule` takes a **natural-language description**, not flags: + ```text -/schedule create "nightly-repro" --cron "0 6 * * *" \ - --prompt "Run /audit-reproducibility against the current passport. If any claim FAILs (not EXPLAINED), summarize which tables are affected and push a notification; otherwise exit quietly." +/schedule nightly at 6am: run /audit-reproducibility against the current passport. +If any claim FAILs (not EXPLAINED), summarize which tables are affected and notify me; +otherwise exit quietly. ``` -`scripts/nightly-repro-check.sh` is a thin local equivalent for users who prefer a machine cron over a Routine (note: a local cron does not survive a closed laptop — prefer `/schedule`). +A precise cron expression (e.g. `0 6 * * *`) is applied via `/schedule update` *after* the routine exists; manage with `/schedule list` / `update` / `remove`. Two scheduling constraints to design around: the **minimum interval is 1 hour**, and accounts carry a **daily run cap** — so batch checks into one routine rather than many small ones. Routines operate on **committed repos**: anything uncommitted or private-by-design (e.g. a local research vault) is invisible to them. + +`scripts/nightly-repro-check.sh` is a thin local equivalent for users who prefer a machine cron over a Routine — and the right tool for uncommitted/private material (note: a local cron does not survive a closed laptop; for committed repos prefer `/schedule`). ## Guardrails for unattended runs - **Never point an unattended loop at a submission portal, shared/restricted data, or a co-author's inbox without a human gate.** Routines *propose*; a human *sends*. (`/triage-inbox` never auto-sends; the [`git-guardrails`](../hooks/git-guardrails.py) hook still blocks destructive git even in a routine.) - **Bound the cost.** A nightly full-manuscript re-audit is fine; a nightly 7× `/seven-pass-review` is not — cost-pilot first. -- **MCP may be absent in headless/cron runs** (interactively-authenticated servers like Gmail). Routines that need them must degrade gracefully. +- **Connectors are INCLUDED by default — least-privilege them.** Cloud Routines run with **all of your claude.ai connectors attached, write access included, and no approval prompts**. An unattended routine that only needs to read your repo should have Gmail/Calendar/Slack *removed from that routine's connector list* before it ever fires — the risk is not a missing connector but a fully-armed one acting without you. (Locally-authenticated MCP servers in your *terminal* sessions are a separate thing and may still be absent in other headless contexts — degrade gracefully either way.) ## Cross-references diff --git a/.claude/references/v2.0-backlog.md b/.claude/references/v2.0-backlog.md index 24b453dcf..d95341f46 100644 --- a/.claude/references/v2.0-backlog.md +++ b/.claude/references/v2.0-backlog.md @@ -36,6 +36,15 @@ The v2.0.0 "modernization" release (2026-06-09) shipped the loop-first / gate-en ## Still deferred (post-v2.0) +### From the 2026-06-10 currency audit (do-soon tier) + +- **Awesome-list distribution** — after topics/homepage/CITATION.cff land: one-line PRs to `hanlulong/awesome-ai-for-economists` (CC0) and the main claude-code awesome list(s). Low effort, high reach for the exact audience. +- **Zenodo DOI** — connect the GitHub repo to Zenodo (owner's account), then the next tagged release mints a DOI automatically; add the DOI badge to README + the `doi` field to CITATION.cff. The citability flywheel for replication-package acknowledgements. +- **Hook touchpoints note** for `orchestrator-protocol.md` / `scheduled-routines.md` — optional wiring (reduce-trigger / failure-resilience / per-reviewer logging hook events). *Verify each event name against current docs before documenting* — this audit caught fabricated CLI syntax; don't repeat it with hook names. +- **Persona-segmented cost table** (grad student vs faculty) — a small "what a month costs" table in the guide's Cost-Conscious Composition section; the model-routing rule now carries a one-line version. +- **Non-Claude coauthor gate note** — document the path for a coauthor who pulls the repo *without* Claude Code: `./scripts/install-hooks.sh` works for them too (pre-commit is plain bash/python), but skills/hooks guidance assumes the Claude loop; a short "for your non-Claude coauthor" README box would close it. +- **Audience scope (owner-set, 2026-06-10):** econ + closely related fields. The psychology/sociology/public-health cards below remain deferred indefinitely — do not build without an explicit owner ask. + ### Portfolio hub: package development — Stata & Python (R shipped in v1.10.0) The R package triad shipped in v1.10.0. Stata and Python remain, to cover the rest of the owner's package portfolio (2 Stata packages, 1 Python package). diff --git a/.claude/rules/model-routing.md b/.claude/rules/model-routing.md index b421623dd..528bb51c5 100644 --- a/.claude/rules/model-routing.md +++ b/.claude/rules/model-routing.md @@ -39,6 +39,17 @@ Model tier is the second cost lever; **effort is the first.** Every model runs a Set per skill/agent with the `effort:` frontmatter field. Several skills ship at `effort: high` for genuinely hard gates (e.g. `/seven-pass-review`, `/simulation-study`, `/r-package-check`). Match effort to the cognitive demand the same way you match model tier — and tune effort before you swap models. +## Where Fable 5 fits — and where it does not + +**Fable 5** (GA 2026-06-09) is the most capable model in Claude Code — and this rule deliberately does **not** route any of the template's fleet to it. Two verified reasons: + +1. **Cost discipline.** Fable 5 is $10/$50 per MTok vs Opus 4.8's $5/$25 — a flat 2× on exactly the judgment tier the 70/20/10 split exists to guard. The referee/editor/verifier agents are bounded, single-sitting tasks; Fable's premium is priced for *long-horizon, larger-than-one-sitting* autonomous work, which the fleet is not. +2. **Protocol maturity.** In launch-week testing, Fable 5 subagents failed the forced structured-output tool protocol 28/28 times where Opus 4.8 was 100% reliable. In a fan-out fleet, a silent tool-protocol failure means a review lens returns *nothing* — the worst failure mode for a review you're trusting. (Same logic as the "don't push Opus down a tier" anti-pattern: a too-immature judge is as bad as a too-cheap one.) + +**Where Fable 5 *is* the right call:** your own interactive sessions on the hardest long-horizon work — a multi-day refactor, a deep research synthesis you'll steer by hand — where you are in the loop to catch a protocol hiccup and the task actually exploits the model's horizon. Select it per-session (`/model fable`); leave the fleet's `model:` pins alone. Re-evaluate at Fable point releases (the protocol gap is the kind of thing that gets fixed); when it does, the high-judgment tier is the natural first candidate. + +**Cost reality check (grad-student budgets):** a full `/review-paper --peer` runs a meaningful fraction of a dollar-denominated token budget at Opus prices; doubling the judgment tier doubles that line item with no quality evidence yet. When cost-constrained, drop *effort* first (the first lever, above), then tier — never the reverse. + ## Why this matters Cost reduction on routed skills is typically **50–80%** with no quality loss on the mechanical tier. The cache-TTL change (5-min default in 2026; Claude subscriptions get 1-hour automatically, API keys opt in) made multi-turn pipelines on API keys materially more expensive; per-agent routing recovers that lost ground without sacrificing the high-judgment lens where it matters. diff --git a/.claude/scripts/statusline.sh b/.claude/scripts/statusline.sh index 4f7ce4789..0ca17df8e 100755 --- a/.claude/scripts/statusline.sh +++ b/.claude/scripts/statusline.sh @@ -2,7 +2,7 @@ # Claude Code status line: shows permission mode, model, and git branch. # # Claude Code pipes a JSON session snapshot to stdin. Relevant keys: -# .model.display_name e.g. "Opus 4.x" +# .model.display_name e.g. "Opus " # .permission_mode e.g. "bypassPermissions" | "plan" | "acceptEdits" | "default" # .workspace.current_dir absolute path of the cwd # diff --git a/.claude/skills/submission-disclosures/SKILL.md b/.claude/skills/submission-disclosures/SKILL.md new file mode 100644 index 000000000..0e7f06681 --- /dev/null +++ b/.claude/skills/submission-disclosures/SKILL.md @@ -0,0 +1,76 @@ +--- +name: submission-disclosures +description: Generate the submission-time disclosure block for a manuscript — the AI-use disclosure statement matched to the target journal's policy, CRediT author-contribution roles, conflict-of-interest statement, and data-availability statement. Use when the user says "AI disclosure", "disclosure statement", "do I need to disclose Claude", "CRediT roles", "conflict of interest statement", "data availability statement", or is preparing a submission package. NOT statistical-disclosure screening of restricted-data outputs — that is /disclosure-check. +argument-hint: "[manuscript path] [journal short-name, e.g. AER] [--no-ai | --statements-only]" +allowed-tools: ["Read", "Grep", "Glob", "Write", "WebSearch", "WebFetch"] +effort: medium +--- + +# /submission-disclosures — The Submission-Time Disclosure Block + +Draft the four statements journals now require (or strongly expect) at submission, in one pass: **AI-use disclosure**, **CRediT contributor roles**, **conflict-of-interest**, and **data availability**. Journals tightened AI-use policies through 2025–2026; an undisclosed-AI finding at a top journal is now a research-integrity problem, not a formatting one — so the statement should be drafted *deliberately*, not improvised in the submission portal at midnight. + +**This skill is about the author's disclosures TO the journal.** It is unrelated to [`/disclosure-check`](../disclosure-check/SKILL.md), which screens restricted-data *outputs* for statistical-disclosure risk (small cells, PII). Same word, different worlds. + +## When to use + +- Preparing a submission or resubmission package and the portal asks for AI-use / COI / data-availability statements. +- A revise-and-resubmit at a journal that adopted an AI policy since the original submission. +- A coauthor asks "do we need to say we used Claude/Copilot/ChatGPT on this?" + +## Phases + +### Phase 1 — Resolve the journal's actual policy + +1. If a journal short-name is given, read its profile in [`journal-profiles.md`](../../references/journal-profiles.md) (top-5 econ + AEA-imprint policy notes + poli-sci top-3). +2. **Verify the current policy on the journal's own site** (`WebSearch`/`WebFetch`: " artificial intelligence policy authors", the journal's submission guidelines page). Policies moved fast in 2025–2026; a cached or remembered policy is not good enough for a submission. Record the URL and retrieval date in the output. +3. If no explicit AI policy exists, default to the strictest common denominator (disclose tools, scope of use, and human responsibility) — over-disclosure is free; under-disclosure is not. + +### Phase 2 — Inventory what was actually used + +Interview briefly (or infer from the repo when evident — e.g. `quality_reports/`, session logs, a CLAUDE.md): + +- **Which tools** (Claude Code, Copilot, ChatGPT, Grammarly-class) and **for what**: writing/editing prose, code authoring, code review, literature search, data analysis, translation. +- **What stayed human**: research design, identification choices, interpretation, final verification of every number and citation (tie to the repo's own verification story — `/audit-reproducibility`, `/verify-claims` — when true, *say so*: "all AI-assisted numbers were independently verified against code" is a strength, not a confession). +- **What AI was NOT used for** when the journal cares (e.g., most policies bar AI as a listed author and bar undisclosed AI-generated images/data). + +### Phase 3 — Draft the four statements + +Write `quality_reports/submission_disclosures_[manuscript-slug].md` containing: + +1. **AI-use disclosure** — journal-matched wording: tools + versions, scope of use, the affirmation that authors take full responsibility for all content and verified all AI-assisted output. Honest and specific; never boilerplate that overclaims ("no AI was used") when the repo's own logs say otherwise. +2. **CRediT roles** — the 14-role taxonomy mapped to each author (interview for the mapping; flag roles no author holds). +3. **Conflict-of-interest** — funding sources, paid/unpaid positions, data-provider relationships (IRB/data-use agreements often constrain what must be stated; cross-ref [`confidential-data.md`](../../rules/confidential-data.md)). +4. **Data availability** — aligned with the replication deposit: openICPSR/DCAS language when the target is an AEA-imprint journal (delegate the deposit itself to [`/replication-package`](../replication-package/SKILL.md); restricted-data access language per [`confidential-data.md`](../../rules/confidential-data.md)). + +With `--statements-only`, emit the statements to chat without writing the file. With `--no-ai`, skip statement 1 (the user asserts no AI assistance — note in chat that the repo's own session logs may contradict this, if they visibly do). + +### Phase 4 — Parity check against the manuscript + +Grep the manuscript for an existing acknowledgments/disclosure section; flag contradictions (e.g., the paper thanks "research assistance" that the COI omits, or an existing AI statement that the new one contradicts). Do not silently overwrite — surface the diff. + +## Exit behavior + +- **Statements drafted, policy verified:** write the file, print the four statements + the policy URL/date, and remind the user the statements are drafts for *author* review — sign-off is theirs. +- **Journal policy unverifiable** (site unreachable, no policy found): emit the strict-default statements, clearly marked "default wording — verify against the journal's current author guidelines before submission." +- **Inventory contradicts `--no-ai`:** stop and surface the contradiction; never produce a false "no AI" statement. + +## Flags + +- `--no-ai` — Skip the AI-use statement (user asserts none was used). The skill still warns if repo evidence visibly contradicts the assertion. +- `--statements-only` — Print the statements to chat; write no file. + +## Cross-references + +- [`.claude/skills/disclosure-check/SKILL.md`](../disclosure-check/SKILL.md) — statistical-disclosure screening of restricted-data outputs (the other "disclosure"; unrelated). +- [`.claude/skills/replication-package/SKILL.md`](../replication-package/SKILL.md) — the deposit the data-availability statement must match. +- [`.claude/references/journal-profiles.md`](../../references/journal-profiles.md) — per-journal calibration, incl. the AEA DCAS policy note. +- [`.claude/rules/confidential-data.md`](../../rules/confidential-data.md) — restricted-data constraints on what the statements can say. +- [`.claude/skills/humanize/SKILL.md`](../humanize/SKILL.md) — detecting AI-voice in prose; disclosure and voice are separate obligations. + +## What this skill does NOT do + +- **Screen outputs for statistical disclosure risk** — that is [`/disclosure-check`](../disclosure-check/SKILL.md). +- **Build the replication deposit** — that is [`/replication-package`](../replication-package/SKILL.md); this skill only writes the statement that points at it. +- **Decide your ethics.** It drafts honest statements from what you report and what the repo shows; whether a use *needed* disclosing under a vague policy is the author's call — the skill defaults to disclosure when in doubt. +- **Submit anything.** Statements go in the user's submission package by the user's hand. diff --git a/.github/CODE_OF_CONDUCT.md b/.github/CODE_OF_CONDUCT.md new file mode 100644 index 000000000..7e80b5fa1 --- /dev/null +++ b/.github/CODE_OF_CONDUCT.md @@ -0,0 +1,7 @@ +# Code of Conduct + +This project follows the [Contributor Covenant v2.1](https://www.contributor-covenant.org/version/2/1/code_of_conduct/). + +**The short version:** be kind, be constructive, assume good faith. Critique ideas and code as hard as you like — the whole template is built on adversarial review — but never people. Harassment, discrimination, and personal attacks are not tolerated. + +**Enforcement:** report conduct issues privately to the maintainer (see the email on [psantanna.com](https://psantanna.com)). Reports are handled confidentially. The maintainer may edit, remove, or reject contributions and ban contributors for behavior that violates this code. diff --git a/.github/CONTRIBUTING.md b/.github/CONTRIBUTING.md index b2fbf19d2..d557c9a7e 100644 --- a/.github/CONTRIBUTING.md +++ b/.github/CONTRIBUTING.md @@ -37,7 +37,7 @@ This repository is a **template** designed for academic researchers to fork and - **Branch naming**: `feat/short-name`, `fix/short-name`, `chore/short-name`, `docs/short-name`. - **Commit messages**: imperative mood ("add", "fix", "refactor"), explain *why* in the body. -- **Co-author Claude** if Claude Code helped: `Co-Authored-By: Claude Opus 4.6 `. +- **Co-author Claude** if Claude Code helped: `Co-Authored-By: Claude ` (version-free — model names drift). - **Use the PR template** (auto-loaded when you open a PR). - **Squash before merging** if your branch has many WIP commits. diff --git a/.github/SECURITY.md b/.github/SECURITY.md new file mode 100644 index 000000000..c94c1cab2 --- /dev/null +++ b/.github/SECURITY.md @@ -0,0 +1,24 @@ +# Security Policy + +This is a template for academic research workflows. It ships **hooks that execute locally** (`.claude/hooks/*.py`, `.githooks/pre-commit`), **skills that drive autonomous review loops**, and **guidance for handling restricted/confidential data** — so security reports are taken seriously even though the repo contains no service or secrets itself. + +## Reporting a vulnerability + +- **Preferred:** open a [private security advisory](https://github.com/pedrohcgs/claude-code-my-workflow/security/advisories/new) on GitHub. +- Please do **not** open a public issue for anything that could expose a forker's data (e.g., a hook that leaks file contents, a guardrail bypass, an unattended-routine footgun). + +In scope, for example: + +- Bypasses of `git-guardrails.py` (destructive-git blocking) or the pre-commit gate. +- A skill/hook that could exfiltrate or overwrite user data outside the repo. +- Incorrect security guidance (e.g., a documented pattern that claims to sandbox but does not — see the `allowed-tools` vs `disallowed-tools` note in `templates/skill-template.md`). +- Flaws in the confidential-data guidance (`.claude/rules/confidential-data.md`) that could lead to restricted-data exposure. + +## What to expect + +Solo-maintained academic project: acknowledgement within a week is the goal, a fix or documented mitigation as soon as practical. Credit given in the CHANGELOG unless you prefer otherwise. + +## Not in scope + +- Vulnerabilities in Claude Code itself → report to [Anthropic](https://www.anthropic.com/responsible-disclosure-policy). +- Vulnerabilities in third-party R/Stata/Python packages the skills orchestrate → report upstream. diff --git a/CHANGELOG.md b/CHANGELOG.md index f18ff44bf..f7536f0a1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,33 @@ If you have forked this template, see the **Upgrading** section at the bottom fo --- +## v2.1.0 — 2026-06-10 + +A **currency + citability release**, driven by a 48-agent web-verified audit ("is this actually up to date and the best for economists, today?"). The architecture audit came back clean — the fixes are facts, not structure. + +### Changed — model refresh (Fable 5) + +- **`model-versions.md` SSoT** — Claude **Fable 5** (GA 2026-06-09) added as the top tier: `claude-fable-5` (alias `fable`), opt-in, $10/$50 per MTok, 1M context, falls back to Opus 4.8 on flagged content, needs Claude Code ≥ 2.1.170. Opus 4.8 demoted from "newest" to "current Opus tier; API/account default". Marker, table, and verified-date updated per the file's own protocol. +- **`model-routing.md`** — new section "Where Fable 5 fits — and where it does not": the fleet deliberately stays on Opus/Sonnet/Haiku (2× price on the judgment tier + launch-week forced-tool-protocol failures observed at 28/28 vs 0 on Opus 4.8); Fable 5 recommended only for interactive long-horizon sessions. Re-evaluate at point releases. +- **`check-model-versions.sh` hardened** — version regex generalized (`4.x` → any `major[.minor]`, so bare "Fable 5" is tracked), Fable added to the tier loop, and a new **superlative-drift check** (flags "newest/most capable" claims naming a non-top tier) — the exact class of claim the Fable launch falsified while the old gate stayed green. SSoT update protocol gains a manual superlative-grep step. +- README model-lineup bullet rewritten (Fable 5 = most capable, Opus 4.8 = default + routed tier); CONTRIBUTING's Co-Authored-By example made version-free. + +### Fixed — Claude Code doc currency (3 real bugs) + +- **`scheduled-routines.md`**: the `/schedule` example used a **fabricated flag syntax** (`--cron/--prompt`) — replaced with the real natural-language form + `/schedule update` for precise cron, the 1-hour minimum interval, the daily run cap, and the committed-repos-only constraint. The MCP guardrail was **inverted**: cloud Routines include **all connectors with write access by default** (the risk is a fully-armed connector, not a missing one) — rewritten to least-privilege guidance. +- **`templates/skill-template.md` + guide**: `allowed-tools` was taught as a restriction; per current docs it is a **pre-approval list** ("does not restrict which tools are available: every tool remains callable"). The actual restrictor is **`disallowed-tools`** — now documented, with a read-only-skill pattern (`disallowed-tools: ["Edit","Write","Bash"]` + `AskUserQuestion` for unattended loops) and new `paths` / `when_to_use` / `arguments` rows. Security note rewritten; guide gains an "`allowed-tools` does not sandbox" warning callout. + +### Added + +- **`/submission-disclosures`** — the submission-time disclosure block: **AI-use disclosure** matched to the target journal's *verified-current* policy (fetched at draft time, not remembered), CRediT roles, COI, and data-availability statements. Two independent audits flagged this as the most economist-salient 2026 norms gap (AEA-family journals now mandate AI-use statements). Distinct from `/disclosure-check` (statistical disclosure). +- **`CITATION.cff`** — the repo is now citable (GitHub "Cite this repository"); Zenodo DOI is a follow-up (backlog). +- **`.github/SECURITY.md` + `.github/CODE_OF_CONDUCT.md`** — community-health files a public template shipping hooks + autonomous loops should have. +- **GitHub discoverability** — repo topics (was: zero) + homepage URL set. + +**Inventory at release: 52 skills, 18 agents, 32 rules, 7 hooks** (was 51 / 18 / 32 / 7 at v2.0.0). + +--- + ## v2.0.0 — 2026-06-09 A **paradigm-shift major release.** The template moves from a *prompt-craft contractor* — craft a prompt, invoke a skill, read a report — to a **verification-gated research lab**: you state a goal, a fleet of specialist agents does the labor under **gates that enforce themselves**, and you act as the **auditor of the disagreements they surface**. Two ideas converge: *"loops, not prompts"* (Boris Cherny / the Claude Code team) and *"ground truth is a process, not a dataset"* (Amazon Science). The result modernizes the **orchestration**, not the substance — the passport, simulation contract, and journal-calibrated referees are untouched. Shipped against `quality_reports/plans/2026-06-09_v2.0-modernization-master-plan.md`. diff --git a/CITATION.cff b/CITATION.cff new file mode 100644 index 000000000..0e35d7ac5 --- /dev/null +++ b/CITATION.cff @@ -0,0 +1,29 @@ +cff-version: 1.2.0 +message: "If this template contributed to your research workflow, please cite it as below." +title: "Claude Code Academic Workflow: a verification-gated research template for economists" +type: software +authors: + - family-names: "Sant'Anna" + given-names: "Pedro H. C." + website: "https://psantanna.com" +repository-code: "https://github.com/pedrohcgs/claude-code-my-workflow" +url: "https://psantanna.com/claude-code-my-workflow/" +license: MIT +version: 2.1.0 +date-released: 2026-06-10 +keywords: + - claude-code + - economics + - econometrics + - reproducible-research + - causal-inference + - difference-in-differences + - academic-writing + - research-workflow +abstract: >- + A ready-to-fork Claude Code template for academic research — slides, papers, + data analysis, and Monte Carlo simulation studies. Verification-gated by + design: fan-out review fleets with hallucination gates, numeric-claim + reproducibility audits against code, an enforcing pre-commit quality gate, + and a validated difference-in-differences / event-study workflow built on + the Callaway–Sant'Anna ecosystem. diff --git a/CLAUDE.md b/CLAUDE.md index befd2de3e..d9ff8c4e2 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -87,7 +87,7 @@ Enforced by `/commit` (halts + asks for override) **and** — once you run `./sc The full table of all skills lives in [README.md](README.md#skills-claudeskills). Most-used, by workflow: - **Slides / teaching:** `/create-lecture` `/compile-latex` `/deploy` `/qa-quarto` `/slide-excellence` `/syllabus` `/teach-from-paper` `/scaffold-exercises` -- **Papers / review:** `/review-paper` (`--peer`) `/seven-pass-review` `/respond-to-referees` `/verify-claims` `/proofread` `/humanize` +- **Papers / review:** `/review-paper` (`--peer`) `/seven-pass-review` `/respond-to-referees` `/verify-claims` `/proofread` `/humanize` `/submission-disclosures` - **Data / reproducibility:** `/data-analysis` `/did-event-study` `/simulation-study` `/audit-reproducibility` `/diagnose` `/replication-package` `/capture-environment` `/power-analysis` `/disclosure-check` - **Research / writing:** `/interview-me` `/lit-review` `/research-ideation` `/preregister` `/grant-proposal` `/data-management-plan` - **Meta / workflow:** `/commit` `/learn` `/new-skill` `/checkpoint` `/context-status` `/deep-audit` `/coauthor-brief` `/triage-inbox` diff --git a/README.md b/README.md index f7109e568..ef041a533 100644 --- a/README.md +++ b/README.md @@ -147,12 +147,12 @@ It covers: The guide covers Claude Code's latest capabilities: -- **Model lineup** — **Opus 4.8** (`claude-opus-4-8`) is the newest model and the API default (GA 2026-05-28, $5/$25 per MTok, 1M context, defaults to `high` effort); Opus 4.7 is the prior generation. Sonnet 4.6 is the workhorse (1M context); Haiku 4.5 the fast tier. Sonnet 4 + original Opus 4 retire 2026-06-15 → migrate to Sonnet 4.6 / Opus 4.8. *(Verified against Anthropic docs 2026-05-31.)* +- **Model lineup** — **Fable 5** (`claude-fable-5`, opt-in via `/model fable` or the `best` alias) is the most capable Claude Code model: Mythos-class, GA 2026-06-09, $10/$50 per MTok, 1M context (128k max output), built for long-horizon agentic work; it falls back to Opus 4.8 on flagged cyber/bio content and needs Claude Code ≥ 2.1.170. **Opus 4.8** (`claude-opus-4-8`) remains the API/account default (GA 2026-05-28, $5/$25 per MTok, 1M context, defaults to `high` effort) — and remains this template's routed high-judgment tier (see `model-routing.md` for why). Sonnet 4.6 is the workhorse (1M context); Haiku 4.5 the fast tier. Sonnet 4 + original Opus 4 retire 2026-06-15 → migrate to Sonnet 4.6 / Opus 4.8. *(Verified against Anthropic docs 2026-06-10.)* - **Effort levels** — `/effort` sets cost vs. thoroughness (`low` / `medium` / `high` / `xhigh` / `max`). **Opus 4.8 defaults to `high`** — its `high` does roughly what 4.7's `xhigh` did for fewer tokens, so reserve `xhigh` for extended exploration and `ultracode` (xhigh + dynamic workflows) for the largest autonomous runs. - **`/goal `** (v1.9.0; Anthropic May 2026) — keep working across turns until a fast model confirms the condition holds. Pairs with `/commit` quality gates for verified-end-state runs. - **`claude agents` dashboard** (v1.9.0; Anthropic May 2026) — single screen for parallel review work (`/review-paper --peer`, `/slide-excellence`). - **Cost-Conscious Composition** — prompt-cache TTL (5-min default on API keys; **1-hour automatic on Claude subscriptions**), 70/20/10 model routing (Haiku/Sonnet/Opus), `/cost` + `/usage` monitoring, Agent SDK credit-pool split (2026-06-15). -- **Skill frontmatter** — `effort`, `context: fork`, `agent`, `hooks`, `disable-model-invocation` (v1.8.0+), and dynamic content (`$ARGUMENTS`, `!command` syntax) +- **Skill frontmatter** — `effort`, `context: fork`, `agent`, `hooks`, `disable-model-invocation` (v1.8.0+), `disallowed-tools` (the *actual* tool restriction — `allowed-tools` only pre-approves), `paths` (glob-scoped auto-activation), and dynamic content (`$ARGUMENTS`, `!command` syntax) - **Permission modes** — Normal, Auto-accept, Plan, Auto (classifier-gated; on Team / Enterprise / API and rolling out to Max; needs Opus 4.6+ or Sonnet 4.6), Bypass - **Hook handler types** — command, prompt, and HTTP handlers with 20+ hook events; hooks see `effort.level` and `$CLAUDE_EFFORT` (Apr 2026 Week 19) - **Advanced agent configuration** — model, maxTurns, isolation, tool restrictions; `model-routing.md` rule codifies per-agent tier (v1.9.0) @@ -188,7 +188,7 @@ This workflow is designed as a **single hub for an entire research program** — ## What's Included
-18 agents, 51 skills, 32 rules, 7 hooks (click to expand) +18 agents, 52 skills, 32 rules, 7 hooks (click to expand) ### Agents (`.claude/agents/`) @@ -265,6 +265,7 @@ This workflow is designed as a **single hub for an entire research program** — | `/coauthor-brief` (v2.0) | Collaborator handoff brief — what changed since last brief, per-artifact state, open questions, reproduce-locally + restricted-data access steps | | `/triage-inbox` (v2.0) | Schedulable academic inbox + calendar triage via Gmail/Calendar MCP — classifies referee requests, R&R/editor, co-author threads, seminar/conference invites, grant/admin deadlines; proposes one human-gated action each (draft reply, calendar hold, `/new-referee-project`, `/coauthor-brief`, snooze); emits a digest + referee-obligations tracker; degrades gracefully when MCP is absent; never auto-sends | | `/diagnose` (v2.0) | Root-cause a wrong/failing empirical result — disciplined reproduce → minimise → hypothesise → instrument → fix loop; tuned for research-code bugs (type coercion, NA/merge blow-ups, clustering/SE choice, seed/package-version drift); `--no-fix` localizes without editing | +| `/submission-disclosures` (v2.1) | The submission-time disclosure block: AI-use disclosure matched to the target journal's verified-current policy, CRediT contributor roles, conflict-of-interest, and data-availability statements (NOT statistical disclosure — that's `/disclosure-check`) | | `/syllabus` (v2.0) | Build/restructure a course syllabus from a topic or reading list — course description + prerequisites, week-by-week schedule (topic→readings→deliverables), measurable learning objectives, assessment scheme + rubric, standard policies (late work / AI use / integrity / accessibility), and a per-week work-list mapping weeks to `/create-lecture` decks; economics-aware (PhD metrics/micro/macro sequences, undergrad) | | `/teach-from-paper` (v2.0) | Reads a paper end-to-end and pitches it to a stated audience level — lecture outline (motivation → setup → key result → method → takeaways), the 3-5 results worth presenting with intuition, a slide skeleton for `/create-lecture`, discussion questions, and a problem-set brief for `/scaffold-exercises` | | `/respond-to-eval` (v2.0) | Teaching analogue of `/respond-to-referees` — clusters course-eval comments into themes, weights by frequency (signal vs noise), classifies Keep / Change / Investigate / Out-of-scope, and drafts concrete changes mapped to the syllabus + slide decks; saves the plan to `quality_reports/teaching/` | diff --git a/docs/index.html b/docs/index.html index 2d812e0d1..760d293f6 100644 --- a/docs/index.html +++ b/docs/index.html @@ -12,7 +12,7 @@ - + @@ -209,7 +209,7 @@

What's in the template

  • 18 specialized agents — proofreader, slide auditor, pedagogy reviewer, R reviewer, TikZ critic, domain reviewer, adversarial QA pair, translator, verifier, claim-verifier (Chain-of-Verification), the simulated-peer-review trio (editor, domain referee, methods referee), the Monte Carlo sim-reviewer, the CRAN-focused R package-reviewer, the AI-voice humanize-auditor, and the promote-memory council — all pinned to a fixed model + effort tier (no longer inherit)
  • Adversarial critic-fixer loop — two agents that check each other's work and loop until dry (converge after 2 consecutive clean rounds; a 5-round cap is only a fallback) — the critic can't fix, the fixer can't approve
  • Quality scoring with mandatory verification — every file scored 0–100; /commit halts below 80 (user can override with an explicit reason); skills that implement the orchestrator pattern verify every output before reporting done
  • -
  • 51 slash commands + 32 context-aware rules + 7 hooks /compile-latex, /proofread, /deploy, /commit, /qa-quarto, /lit-review, /review-paper, /respond-to-referees, /new-diagram, /data-analysis, /simulation-study, /r-package-check, /audit-reproducibility, /checkpoint, /preregister, /replication-package, /disclosure-check, /grant-proposal, /syllabus, /triage-inbox, and more — plus quality gates, TikZ prevention/measurement, notation consistency, R conventions, and the ambient prompt-shaping rule (no /prompt command)
  • +
  • 52 slash commands + 32 context-aware rules + 7 hooks /compile-latex, /proofread, /deploy, /commit, /qa-quarto, /lit-review, /review-paper, /respond-to-referees, /new-diagram, /data-analysis, /simulation-study, /r-package-check, /audit-reproducibility, /checkpoint, /preregister, /replication-package, /disclosure-check, /grant-proposal, /syllabus, /triage-inbox, and more — plus quality gates, TikZ prevention/measurement, notation consistency, R conventions, and the ambient prompt-shaping rule (no /prompt command)
  • Simulated peer review /review-paper --peer <journal> runs a full editorial pipeline: editor desk review, two blind referees with deliberately different dispositions (STRUCTURAL / CREDIBILITY / MEASUREMENT / POLICY / THEORY / SKEPTIC), editorial synthesis with FATAL / ADDRESSABLE / TASTE classification. Calibrated to 5 econ journals (AER / QJE / JPE / ECMA / ReStud) plus a template for adding your own field. Adapted from Hugo Sant’Anna’s clo-author with permission.
  • Research workflow skills /lit-review for literature synthesis, /research-ideation for hypothesis generation, /interview-me to formalize ideas, /review-paper for manuscript review, /data-analysis for end-to-end R analysis
  • Smart hooks (7) — desktop notifications (macOS/Linux); pre-compaction context snapshots to session logs (DRAFT-block default ON); progressive context-usage warnings; log-reminder.py auto-writes the session log on every meaningful change-set; git-guardrails.py blocks dangerous git ops (reset --hard, clean -f, push --force, add -A); claim-reconcile.py flags stale numeric claims when scripts change; plus a real git pre-commit hook (.githooks/pre-commit via ./scripts/install-hooks.sh) that runs the surface-sync + quality gates
  • diff --git a/docs/workflow-guide.html b/docs/workflow-guide.html index 630538711..ed6c4495e 100644 --- a/docs/workflow-guide.html +++ b/docs/workflow-guide.html @@ -2558,7 +2558,7 @@

    My Claude Code Setup

    Modified
    -

    June 9, 2026

    +

    June 10, 2026

    @@ -2757,7 +2757,7 @@

    -

    You talk, Claude orchestrates. The 18 agents, 51 skills, and 32 rules exist so you don’t have to think about them. Describe your goal, approve the plan, and let the system work.

    +

    You talk, Claude orchestrates. The 18 agents, 52 skills, and 32 rules exist so you don’t have to think about them. Describe your goal, approve the plan, and let the system work.

    @@ -2770,7 +2770,7 @@

    -

    This guide describes the full system — 18 agents, 51 skills, 32 rules. That is the ceiling, not the floor. Start with just CLAUDE.md and 2–3 skills (/compile-latex, /proofread, /commit). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you’re ready.

    +

    This guide describes the full system — 18 agents, 52 skills, 32 rules. That is the ceiling, not the floor. Start with just CLAUDE.md and 2–3 skills (/compile-latex, /proofread, /commit). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you’re ready.


    @@ -3597,7 +3597,7 @@

    -

    Claude Code ships with built-in skills beyond this template’s 51: /batch orchestrates parallel refactoring across your codebase (using git worktrees for isolation), /simplify runs 3-agent code review and applies fixes, and /debug helps troubleshoot sessions. These complement the academic skills above.

    +

    Claude Code ships with built-in skills beyond this template’s 52: /batch orchestrates parallel refactoring across your codebase (using git worktrees for isolation), /simplify runs 3-agent code review and applies fixes, and /debug helps troubleshoot sessions. These complement the academic skills above.

    @@ -6269,7 +6269,7 @@

    7.4 Step 4: Creating Custom Skills

    -

    The guide includes 51 skills for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own.

    +

    The guide includes 52 skills for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own.

    7.4.1 When to Create a Skill

    Create a skill when: - You repeatedly explain the same 3+ step workflow to Claude - You need domain-specific quality checks (citation style, notation consistency, lab protocols) - You enforce field-specific output formats (thesis structure, journal templates) - You coordinate multi-tool workflows (data → analysis → manuscript)

    @@ -6302,7 +6302,7 @@

    7.4.3 Complete Frontmatter Reference

    -

    The YAML frontmatter controls how your skill behaves. Here are all available fields:

    +

    The YAML frontmatter controls how your skill behaves. The most-used fields:

    @@ -6334,10 +6334,30 @@

    - + + + + + + + + + + + + + + + + + + + + + @@ -6374,6 +6394,19 @@

    allowed-toolsRestrict which tools the skill can usePre-approve tools (skip permission prompts). NOT a restriction — unlisted tools remain callable through normal permissions ["Read", "Bash", "Glob"]
    disallowed-toolsRemove tools from Claude’s pool while the skill is active — this is the actual restriction mechanism["Edit", "Write"]
    pathsGlob patterns scoping when the skill auto-activates (skill-level analogue of path-scoped rules)["scripts/**/*.R"]
    when_to_useExtra routing context for auto-invocationafter a merge conflict
    argumentsNamed positional arguments for $name substitution[infile, outfile]
    effort Override reasoning effort level high
    +
    +
    +
    + +
    +
    +Warningallowed-tools does not sandbox +
    +
    +
    +

    A common (and security-relevant) misreading: omitting Bash from allowed-tools does not make a skill read-only — allowed-tools only pre-approves the listed tools so they skip permission prompts; everything else stays callable under your normal permission settings. To genuinely remove tools while a skill runs, use disallowed-tools (e.g. ["Edit", "Write", "Bash"] for a read-only audit skill, plus AskUserQuestion for anything that runs unattended). See templates/skill-template.md for the full pattern.

    +
    +
    @@ -6558,7 +6591,7 @@

    - -
    +

    For capabilities beyond file editing and shell commands — web search during literature review, database queries for replication, or reference manager integration (Zotero, Mendeley) — Claude Code supports MCP servers. Configure them in .claude/settings.json under "mcpServers". Start with skills and agents first; add MCP when you need external integrations.

    diff --git a/guide/workflow-guide.html b/guide/workflow-guide.html index 630538711..ed6c4495e 100644 --- a/guide/workflow-guide.html +++ b/guide/workflow-guide.html @@ -2558,7 +2558,7 @@

    My Claude Code Setup

    Modified
    -

    June 9, 2026

    +

    June 10, 2026

    @@ -2757,7 +2757,7 @@

    -

    You talk, Claude orchestrates. The 18 agents, 51 skills, and 32 rules exist so you don’t have to think about them. Describe your goal, approve the plan, and let the system work.

    +

    You talk, Claude orchestrates. The 18 agents, 52 skills, and 32 rules exist so you don’t have to think about them. Describe your goal, approve the plan, and let the system work.

    @@ -2770,7 +2770,7 @@

    -

    This guide describes the full system — 18 agents, 51 skills, 32 rules. That is the ceiling, not the floor. Start with just CLAUDE.md and 2–3 skills (/compile-latex, /proofread, /commit). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you’re ready.

    +

    This guide describes the full system — 18 agents, 52 skills, 32 rules. That is the ceiling, not the floor. Start with just CLAUDE.md and 2–3 skills (/compile-latex, /proofread, /commit). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you’re ready.


    @@ -3597,7 +3597,7 @@

    -

    Claude Code ships with built-in skills beyond this template’s 51: /batch orchestrates parallel refactoring across your codebase (using git worktrees for isolation), /simplify runs 3-agent code review and applies fixes, and /debug helps troubleshoot sessions. These complement the academic skills above.

    +

    Claude Code ships with built-in skills beyond this template’s 52: /batch orchestrates parallel refactoring across your codebase (using git worktrees for isolation), /simplify runs 3-agent code review and applies fixes, and /debug helps troubleshoot sessions. These complement the academic skills above.

    @@ -6269,7 +6269,7 @@

    7.4 Step 4: Creating Custom Skills

    -

    The guide includes 51 skills for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own.

    +

    The guide includes 52 skills for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own.

    7.4.1 When to Create a Skill

    Create a skill when: - You repeatedly explain the same 3+ step workflow to Claude - You need domain-specific quality checks (citation style, notation consistency, lab protocols) - You enforce field-specific output formats (thesis structure, journal templates) - You coordinate multi-tool workflows (data → analysis → manuscript)

    @@ -6302,7 +6302,7 @@

    7.4.3 Complete Frontmatter Reference

    -

    The YAML frontmatter controls how your skill behaves. Here are all available fields:

    +

    The YAML frontmatter controls how your skill behaves. The most-used fields:

    @@ -6334,10 +6334,30 @@

    - + + + + + + + + + + + + + + + + + + + + + @@ -6374,6 +6394,19 @@

    allowed-toolsRestrict which tools the skill can usePre-approve tools (skip permission prompts). NOT a restriction — unlisted tools remain callable through normal permissions ["Read", "Bash", "Glob"]
    disallowed-toolsRemove tools from Claude’s pool while the skill is active — this is the actual restriction mechanism["Edit", "Write"]
    pathsGlob patterns scoping when the skill auto-activates (skill-level analogue of path-scoped rules)["scripts/**/*.R"]
    when_to_useExtra routing context for auto-invocationafter a merge conflict
    argumentsNamed positional arguments for $name substitution[infile, outfile]
    effort Override reasoning effort level high
    +
    +
    +
    + +
    +
    +Warningallowed-tools does not sandbox +
    +
    +
    +

    A common (and security-relevant) misreading: omitting Bash from allowed-tools does not make a skill read-only — allowed-tools only pre-approves the listed tools so they skip permission prompts; everything else stays callable under your normal permission settings. To genuinely remove tools while a skill runs, use disallowed-tools (e.g. ["Edit", "Write", "Bash"] for a read-only audit skill, plus AskUserQuestion for anything that runs unattended). See templates/skill-template.md for the full pattern.

    +
    +
    @@ -6558,7 +6591,7 @@

    - -
    +

    For capabilities beyond file editing and shell commands — web search during literature review, database queries for replication, or reference manager integration (Zotero, Mendeley) — Claude Code supports MCP servers. Configure them in .claude/settings.json under "mcpServers". Start with skills and agents first; add MCP when you need external integrations.

    diff --git a/guide/workflow-guide.qmd b/guide/workflow-guide.qmd index 1f3b56894..a499b3565 100644 --- a/guide/workflow-guide.qmd +++ b/guide/workflow-guide.qmd @@ -155,13 +155,13 @@ Most of the time, you just describe what you want and Claude handles the rest. E ::: {.callout-important} ## The Bottom Line -**You talk, Claude orchestrates.** The 18 agents, 51 skills, and 32 rules exist so you don't have to think about them. Describe your goal, approve the plan, and let the system work. +**You talk, Claude orchestrates.** The 18 agents, 52 skills, and 32 rules exist so you don't have to think about them. Describe your goal, approve the plan, and let the system work. ::: ::: {.callout-note} ## You Don't Need All of This on Day One -This guide describes the full system --- 18 agents, 51 skills, 32 rules. That is the ceiling, not the floor. **Start with just CLAUDE.md and 2--3 skills** (`/compile-latex`, `/proofread`, `/commit`). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you're ready. +This guide describes the full system --- 18 agents, 52 skills, 32 rules. That is the ceiling, not the floor. **Start with just CLAUDE.md and 2--3 skills** (`/compile-latex`, `/proofread`, `/commit`). Add rules and agents as you discover what you need. The template is designed for progressive adoption: fork it, fill in the placeholders, and start working. Everything else is there when you're ready. ::: --- @@ -696,7 +696,7 @@ argument-hint: "[filename without .tex extension]" ::: {.callout-note} ## Built-In Skills -Claude Code ships with built-in skills beyond this template's 51: `/batch` orchestrates parallel refactoring across your codebase (using git worktrees for isolation), `/simplify` runs 3-agent code review and applies fixes, and `/debug` helps troubleshoot sessions. These complement the academic skills above. +Claude Code ships with built-in skills beyond this template's 52: `/batch` orchestrates parallel refactoring across your codebase (using git worktrees for isolation), `/simplify` runs 3-agent code review and applies fixes, and `/debug` helps troubleshoot sessions. These complement the academic skills above. ::: ## Agents --- Specialized Reviewers {#agents---specialized-reviewers} @@ -2388,7 +2388,7 @@ The template includes matching LaTeX and Quarto palettes. To customize: ## Step 4: Creating Custom Skills {#sec-create-skills} -The guide includes 51 [skills](#skills---reusable-slash-commands) for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own. +The guide includes 52 [skills](#skills---reusable-slash-commands) for common academic tasks. But if you have repetitive workflows specific to your domain, you can create your own. ### When to Create a Skill @@ -2433,14 +2433,18 @@ Solution: [How to fix] ### Complete Frontmatter Reference {#skill-frontmatter} -The YAML frontmatter controls how your skill behaves. Here are all available fields: +The YAML frontmatter controls how your skill behaves. The most-used fields: | Field | Purpose | Example | |-------|---------|---------| | `name` | Display name in `/` menu | `compile-latex` | | `description` | **Most important.** Controls when Claude auto-loads the skill | `Compile Beamer slides...` | | `argument-hint` | Placeholder shown after `/skill-name` | `` | -| `allowed-tools` | Restrict which tools the skill can use | `["Read", "Bash", "Glob"]` | +| `allowed-tools` | **Pre-approve** tools (skip permission prompts). NOT a restriction — unlisted tools remain callable through normal permissions | `["Read", "Bash", "Glob"]` | +| `disallowed-tools` | **Remove** tools from Claude's pool while the skill is active — this is the actual restriction mechanism | `["Edit", "Write"]` | +| `paths` | Glob patterns scoping when the skill auto-activates (skill-level analogue of path-scoped rules) | `["scripts/**/*.R"]` | +| `when_to_use` | Extra routing context for auto-invocation | `after a merge conflict` | +| `arguments` | Named positional arguments for `$name` substitution | `[infile, outfile]` | | `effort` | Override reasoning effort level | `high` | | `context` | Set to `fork` to run in an isolated subagent | `fork` | | `agent` | Link to an agent definition in `.claude/agents/` | `proofreader` | @@ -2449,6 +2453,12 @@ The YAML frontmatter controls how your skill behaves. Here are all available fie | `disable-model-invocation` | Prevent Claude from auto-triggering | `true` | | `user-invocable` | Whether it appears in the `/` menu | `true` (default) | +::: {.callout-warning} +## `allowed-tools` does not sandbox + +A common (and security-relevant) misreading: omitting `Bash` from `allowed-tools` does **not** make a skill read-only — `allowed-tools` only *pre-approves* the listed tools so they skip permission prompts; everything else stays callable under your normal permission settings. To genuinely remove tools while a skill runs, use `disallowed-tools` (e.g. `["Edit", "Write", "Bash"]` for a read-only audit skill, plus `AskUserQuestion` for anything that runs unattended). See `templates/skill-template.md` for the full pattern. +::: + ::: {.callout-note} ## Key Design Choices diff --git a/scripts/check-model-versions.sh b/scripts/check-model-versions.sh index c737fae4f..b82570e1e 100755 --- a/scripts/check-model-versions.sh +++ b/scripts/check-model-versions.sh @@ -43,24 +43,48 @@ SURFACES=( # A line is allowed to name an older version if it carries one of these markers. ALLOW='prior generation|prior gen|prior Opus|retire|migrat|historical|deprecat|was:|was |or later|incl\. 4\.|rolling out|GA 2026-0|beta|4\.[0-9]+.s |model-allow' +# Version token: "4.8", "5", "5.1" — Fable has no minor version at launch, so +# the regex must accept a bare major (the old `4\.[0-9]+` silently skipped it). +VER='[0-9]+(\.[0-9]+)?' + drift=0 -for tier in "Opus" "Sonnet" "Haiku"; do - current="$(echo "$CURRENT_LINE" | grep -oE "$tier 4\.[0-9]+" | head -1)" +for tier in "Fable" "Opus" "Sonnet" "Haiku"; do + current="$(echo "$CURRENT_LINE" | grep -oE "$tier $VER" | head -1)" [ -n "$current" ] || continue for f in "${SURFACES[@]}"; do [ -f "$REPO/$f" ] || continue while IFS=: read -r lineno text; do - ver="$(echo "$text" | grep -oE "$tier 4\.[0-9]+" | head -1)" + ver="$(echo "$text" | grep -oE "$tier $VER" | head -1)" [ -n "$ver" ] || continue [ "$ver" = "$current" ] && continue # names the current version → fine echo "$text" | grep -qiE "$ALLOW" && continue # allow-marked line → fine echo " $f:$lineno presents '$ver' (current $tier is '$current')" >&2 echo " → $(echo "$text" | sed -E 's/^[[:space:]]+//' | cut -c1-110)" >&2 drift=1 - done < <(grep -nE "$tier 4\.[0-9]+" "$REPO/$f") + done < <(grep -nE "$tier $VER" "$REPO/$f") done done +# Superlative drift: "newest model" / "most capable" claims are SEMANTIC, not +# version strings — the 2026-06-09 Fable 5 launch made "Opus 4.8 is the newest" +# false while the version check above stayed green. Flag any superlative line +# that names a non-top tier (Opus/Sonnet/Haiku) without mentioning the top tier +# (Fable) and without an allow-marker. Tier-relative phrasings ("the newest +# Opus") are fine and excluded. +TOP_TIER="$(echo "$CURRENT_LINE" | sed -E 's/.*` table** *and* keeping the prose counts (e.g. "51 skills") in sync across `README.md`, `docs/index.html`, the guide, and `templates/skill-template.md`. `./scripts/check-surface-sync.sh` enforces **both** the counts and the table rows — run it before you open a PR. +7. **Keep the surfaces in sync** when adding features. Adding a skill means **adding its row to the README `` table** *and* keeping the prose counts (the "NN skills / NN rules" phrasings) in sync across `README.md`, `docs/index.html`, the guide, and `templates/skill-template.md`. `./scripts/check-surface-sync.sh` enforces **both** the counts and the table rows — run it before you open a PR. ## PR style diff --git a/docs/workflow-guide.html b/docs/workflow-guide.html index ed6c4495e..bf0445df8 100644 --- a/docs/workflow-guide.html +++ b/docs/workflow-guide.html @@ -3664,11 +3664,11 @@

    -NoteCurrent Anthropic lineup (verified 2026-05-31) +NoteCurrent Anthropic lineup (verified 2026-06-10)
    -

    Opus 4.8 (claude-opus-4-8) is the newest model and the API default (GA 2026-05-28, $5/$25 per MTok input/output, 1M context); it defaults to high effort. Opus 4.7 is the prior generation. Sonnet 4.6 is the workhorse mid-tier (1M context). Haiku 4.5 is the cost-efficient fast model. Verified against Anthropic docs 2026-05-31.

    +

    Fable 5 (claude-fable-5, alias fable; opt-in via /model fable or the best alias) is the most capable Claude Code model — Mythos-class, GA 2026-06-09, $10/$50 per MTok, 1M context (128k max output), built for long-horizon agentic work; needs Claude Code ≥ 2.1.170 and falls back to Opus 4.8 on flagged content. Opus 4.8 (claude-opus-4-8) is the API/account default and this template’s routed high-judgment tier (GA 2026-05-28, $5/$25 per MTok, 1M context, defaults to high effort) — see the routing rule for why the fleet deliberately stays on it. Opus 4.7 is the prior generation. Sonnet 4.6 is the workhorse mid-tier (1M context). Haiku 4.5 is the cost-efficient fast model. Verified against Anthropic docs 2026-06-10.

    Retirement notice: Sonnet 4 and the original Opus 4 retire on 2026-06-15. If your environment pins one of these (ANTHROPIC_MODEL env, hard-coded model IDs in CI, etc.), migrate before then. See TROUBLESHOOTING.md for the migration checklist.

    diff --git a/guide/workflow-guide.html b/guide/workflow-guide.html index ed6c4495e..bf0445df8 100644 --- a/guide/workflow-guide.html +++ b/guide/workflow-guide.html @@ -3664,11 +3664,11 @@

    -NoteCurrent Anthropic lineup (verified 2026-05-31) +NoteCurrent Anthropic lineup (verified 2026-06-10)
    -

    Opus 4.8 (claude-opus-4-8) is the newest model and the API default (GA 2026-05-28, $5/$25 per MTok input/output, 1M context); it defaults to high effort. Opus 4.7 is the prior generation. Sonnet 4.6 is the workhorse mid-tier (1M context). Haiku 4.5 is the cost-efficient fast model. Verified against Anthropic docs 2026-05-31.

    +

    Fable 5 (claude-fable-5, alias fable; opt-in via /model fable or the best alias) is the most capable Claude Code model — Mythos-class, GA 2026-06-09, $10/$50 per MTok, 1M context (128k max output), built for long-horizon agentic work; needs Claude Code ≥ 2.1.170 and falls back to Opus 4.8 on flagged content. Opus 4.8 (claude-opus-4-8) is the API/account default and this template’s routed high-judgment tier (GA 2026-05-28, $5/$25 per MTok, 1M context, defaults to high effort) — see the routing rule for why the fleet deliberately stays on it. Opus 4.7 is the prior generation. Sonnet 4.6 is the workhorse mid-tier (1M context). Haiku 4.5 is the cost-efficient fast model. Verified against Anthropic docs 2026-06-10.

    Retirement notice: Sonnet 4 and the original Opus 4 retire on 2026-06-15. If your environment pins one of these (ANTHROPIC_MODEL env, hard-coded model IDs in CI, etc.), migrate before then. See TROUBLESHOOTING.md for the migration checklist.

    diff --git a/guide/workflow-guide.qmd b/guide/workflow-guide.qmd index a499b3565..68cfdf67b 100644 --- a/guide/workflow-guide.qmd +++ b/guide/workflow-guide.qmd @@ -758,9 +758,9 @@ Claude Code also offers experimental **Agent Teams** --- multiple independent se ### Multi-Model Strategy: Cost vs. Quality {#multi-model-strategy-cost-vs-quality} ::: {.callout-note} -## Current Anthropic lineup (verified 2026-05-31) +## Current Anthropic lineup (verified 2026-06-10) -**Opus 4.8** (`claude-opus-4-8`) is the newest model and the API default (GA 2026-05-28, \$5/\$25 per MTok input/output, 1M context); it defaults to `high` effort. **Opus 4.7** is the prior generation. **Sonnet 4.6** is the workhorse mid-tier (1M context). **Haiku 4.5** is the cost-efficient fast model. *Verified against Anthropic docs 2026-05-31.* +**Fable 5** (`claude-fable-5`, alias `fable`; opt-in via `/model fable` or the `best` alias) is the most capable Claude Code model — Mythos-class, GA 2026-06-09, \$10/\$50 per MTok, 1M context (128k max output), built for long-horizon agentic work; needs Claude Code ≥ 2.1.170 and falls back to Opus 4.8 on flagged content. **Opus 4.8** (`claude-opus-4-8`) is the API/account default and this template's routed high-judgment tier (GA 2026-05-28, \$5/\$25 per MTok, 1M context, defaults to `high` effort) — see the routing rule for why the fleet deliberately stays on it. **Opus 4.7** is the prior generation. **Sonnet 4.6** is the workhorse mid-tier (1M context). **Haiku 4.5** is the cost-efficient fast model. *Verified against Anthropic docs 2026-06-10.* **Retirement notice:** Sonnet 4 and the original Opus 4 retire on **2026-06-15**. If your environment pins one of these (`ANTHROPIC_MODEL` env, hard-coded model IDs in CI, etc.), migrate before then. See [TROUBLESHOOTING.md](../TROUBLESHOOTING.md) for the migration checklist. ::: diff --git a/scripts/check-model-versions.sh b/scripts/check-model-versions.sh index b82570e1e..034eb8476 100755 --- a/scripts/check-model-versions.sh +++ b/scripts/check-model-versions.sh @@ -78,7 +78,13 @@ for f in "${SURFACES[@]}"; do echo "$text" | grep -qiE "$TOP_TIER" && continue # already credits the top tier echo "$text" | grep -qiE "newest (Opus|Sonnet|Haiku)" && continue # tier-relative superlative → fine echo "$text" | grep -qiE "(Opus|Sonnet|Haiku) $VER" || continue # only flag lines naming a versioned tier - echo "$text" | grep -qiE "$ALLOW" && continue + # NOTE: deliberately NOT short-circuiting on the general $ALLOW list here. + # A superlative is a claim about the WHOLE lineup; an allow-marker earned by a + # different clause in the same sentence ("Opus 4.7 is the prior generation", + # a "GA 2026-.." date) must not suppress it — that exact interaction let + # "Opus 4.8 is the newest model" slip past this check in v2.1 review. The only + # explicit escape is an inline model-allow comment placed for THIS claim. + echo "$text" | grep -q "model-allow" && continue echo " $f:$lineno superlative claim may be stale (top tier is now '$TOP_TIER'):" >&2 echo " → $(echo "$text" | sed -E 's/^[[:space:]]+//' | cut -c1-110)" >&2 drift=1