Deterministic verifier + focused LLM verification protocol for the visual-explainer skill. Target executors: Codex, Opus, Sonnet-tier — the checker must be runnable with ONE command and produce machine-readable results a weak model can act on without judgment.
plugins/visual-explainer/scripts/verify/
ve-verify.mjs # CLI entry
lib/engine.mjs # runs checks, assembles report
lib/profile.mjs # profile auto-detection (page|slides|magazine|poster|video-comp)
lib/static-text.mjs # regex-stage runner
lib/static-dom.mjs # linkedom-stage runner
lib/browser.mjs # playwright-core stage: viewports × schemes, screenshots, probes
checks.json # THE CATALOG — data-driven check definitions (id, stage, profiles,
# applies_when, severity, params). Complex checks reference a named
# impl in lib/checks/<stage>.mjs by id.
rubrics/ # LLM-pass rubric files (one per pass, markdown, ≤ 1 screen each)
pass-hierarchy.md
pass-aesthetic-<preset>.md # one per preset; verifier loads ONLY the active preset's file
pass-completeness.md
pass-artifact-slop-gap.md # explicit semantic gap only; Impeccable/Unslop stay external
evals/
fixtures/violations/<check-id>.html # exactly one seeded violation each
fixtures/clean/<profile>.html # must pass everything (false-positive guard)
run.mjs # runs ve-verify over all fixtures, asserts expectations
expectations.json # fixture -> {must_fire: [ids], must_not_fire: "*"}
node plugins/visual-explainer/scripts/verify/ve-verify.mjs <file.html> \
[--profile page|slides|magazine|poster|video-comp] # default: auto-detect
[--preset mono-industrial|nothing|...] # default: auto-detect
[--json <out.json>] [--screens <dir>] [--static-only] [--quiet]
- Exit 0 = no errors (warns allowed). Exit 1 = ≥1 error-severity failure. Exit 2 = engine crash.
- The verifier runs via bare
nodewith directimport('playwright-core')— no npm/npx at runtime, so it works in minimal agent environments. - playwright-core + linkedom added to package.json devDependencies (pin what's in node_modules: playwright-core 1.60.0).
- Browser stage matrix: {1440×900, 390×844} × {light, dark} → 4 runs; screenshots saved as
<screens>/<viewport>-<scheme>.png+ full-page variants. Console messages + failed requests collected per run. Skip matrix rows per profile (poster: native canvas only; video-comp: 1920×1080 or 1080×1920). - Report JSON:
{
"file": "...", "profile": "page", "preset": "mono-industrial",
"summary": { "errors": 2, "warns": 1, "skipped": 14, "passed": 61 },
"checks": [ { "id": "...", "stage": "...", "severity": "error",
"status": "pass|fail|warn|skip", "evidence": "...",
"where": "selector/line/screenshot ref", "fix_hint": "..." } ],
"screenshots": ["..."],
"llm_passes_required": ["hierarchy", "aesthetic-mono-industrial", "completeness", "impeccable:critique", "unslop:cleanup-report", "artifacture:slop-gap"],
"llm_dispatch_plan": [
{"pass":"hierarchy","owner":"artifacture","status":"ready","model":"provider-small-vision","batch_size":2},
{"pass":"impeccable:critique","owner":"impeccable","status":"delegate-to-installed-skill"}
]
}llm_dispatch_plan is the executable routing boundary. Artifacture-owned
passes are ready only when the resolved eval policy contains a qualified
route; otherwise they are explicitly skipped with
reason: "no-eval-qualified-model". A host must not replace a skipped route
with its current/main-thread model. Companion-skill entries retain their own
skill and model policies.
- Human output: compact table, failures first, each with fix_hint quoting the doc rule.
- slides:
scroll-snap-type: y mandatoryon a track + slides ≥ 100dvh - magazine:
scroll-snap-type: x mandatory - poster: single fixed
w-[Npx] h-[Npx]root or--postermarker / TSX source - video-comp: hyperframes markers (gsap timeline, data-hf attrs, scene structure)
- page: fallback
- preset: explicit
data-ve-presetattribute first (or--presetflag), else STRONG token signatures only (Doto in a font-family declaration = nothing; >=2 core mono token names = mono-industrial); generic font pairings and filenames must never gate preset conformance — "custom" otherwise; loose heuristics may only route LLM passes applies_whenguards: each check declares a detection predicate (regex/selector) evaluated before running; non-applicable → status "skip" with reason. THIS IS THE FALSE-POSITIVE FIREWALL.
- For every
fixtures/violations/<id>.html: run ve-verify (static stages always; browser stage only for browser-stage fixtures), assert<id>fires with expected severity AND no other error-severity check fires (warns from unrelated checks tolerated but reported). - For every
fixtures/clean/<profile>.html: assert zero fires (errors AND warns). - Summary: per-check catch-rate table; exit non-zero if any expectation unmet.
- Runtime budget: static fixtures < 5s total; browser fixtures batched in one chromium instance, < 90s total. This is the regression suite for the checker itself.
Sequenced so a weak model cannot skip or blend steps:
- Run ve-verify. If exit 1 → fix root cause → re-run. Max 3 cycles, then deliver with explicit failure disclosure. (Deterministic gate FIRST; no LLM review of pages that fail mechanics.)
- LLM passes and delegated skills, each a separate small context consuming ONLY its named inputs:
- P1 hierarchy/layout: 4 screenshots + pass-hierarchy.md (squint, weight, moment-of-surprise)
- P2 aesthetic fidelity: 4 screenshots + the ACTIVE preset rubric only (swap test, motif rules)
- P3 completeness: source material inventory + extracted section list/headings (NO screenshots)
- D1 visual craft: candidate screenshots routed to installed Impeccable critique/audit
- D2 prose: excluded-filtered prose routed to Unslop
cleanup --report - P4 artifact slop gap: an explicitly nominated crop plus truth excerpt, limited to false
sequence, state/confidence, or provenance
Claude Code: parallelize compatible Artifacture
ve-verifier-*agents. Invoke Impeccable and Unslop through their installed skills. Codex/single-agent: keep every pass/skill in a fresh context. Artifacture rubric files are authoritative only for Artifacture-owned paths.
- Verdict merge: any local or delegated check fails → fix → re-run only that path. Max 2 cycles.
- Delivery message MUST include: report path, pass/fail per LLM pass, and either "verified" or the explicit could-not-verify disclosure. (Protocol requirement, reviewable from transcript.)
- SKILL.md §6 rewritten (shorter than today: point at ve-verify + verification.md).
- references/verification.md: full protocol incl. rubrics index + report schema.
- references/*.md: add rule-ID anchors where checks cite docs (only where curation demands).
- .claude/agents/ve-verifier-{layout,aesthetic,completeness,artifact-slop-gap}.md
- check.mjs: keep as build smoke test; add
ve:verify+ve:evalnpm scripts (documented as direct-node invocations too, given broken npm shim).
- Fable: freeze checks.json catalog + this SPEC (after curation merge).
- Codex run A: engine + CLI + browser stage. Codex run B (parallel, disjoint files): fixtures + evals/run.mjs. Codex run C (after A/B interfaces exist): docs + rubrics + verifier agents.
- Opus adversarial verify: break the checker (seed tricky false-pos/false-neg pages), audit spec fidelity against docs, review rubric quality for weak-model followability.
- Fable: run evals, final fidelity check, iterate 2–4 until green + no confirmed findings.