All notable changes to @testsprite/testsprite-cli are documented here. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
0.1.1 - 2026-06-12
- README: point the launch video at the updated public asset. Docs-only release — no code changes.
0.1.0 - 2026-06-10
-
testsprite init— one-shot onboarding command that chainsauth configure→auth whoami→agent installin a single interactive invocation. Accepts--from-env,--yes, and--agent <target>for non-interactive and CI use. -
agent install/agent list— write a ready-made TestSprite verification-loop skill file into your project so your coding agent knows the commands, the exit codes, and the failure-bundle layout. Pure-local command: no network, no credentials. Supported targets:claude(GA),codex,cursor,cline,antigravity(experimental). Thecodextarget uses managed-section mode that writes a sentinel-delimited block insideAGENTS.mdwithout clobbering surrounding content.--forcebacks up existing own-file targets before overwriting. -
auth configure/auth whoami/auth logout— API-key management.--from-envreadsTESTSPRITE_API_KEYfor non-interactive setup. Credentials stored at~/.testsprite/credentials(INI, mode0600). -
project list/project get— cursor-paginated project listing and single-project lookup. -
test list/test get— cursor-paginated test listing under a project (with--status,--type,--created-fromfilters) and single-test lookup. -
test create— create a frontend or backend test. Backend tests supply a code file directly (--code-file); frontend tests use--code-fileor generate from a plan-steps document (--plan-from). The--run --waitflags chain create → trigger → poll in one invocation. Dependency metadata flags for backend tests:--produces <var>(repeatable),--needs <var>(repeatable),--category <str>. -
test create-batch— bulk-create frontend tests from a JSONL plan file (--plans) or a directory of plan files (--plan-from-dir). Optional--run --max-concurrency <N>fans out triggers. -
test update <test-id>/test delete <test-id>/test delete-batch— metadata update (name, description) and permanent hard-delete of one or many tests.--confirmis required for destructive operations.test delete-batchsupports--all --project <id>and--status <list>for bulk targeted deletes. -
test code get <test-id>/test code put <test-id>— read the generated test source and replace it with etag-guarded optimistic concurrency (--expected-version, or--forceto skip the guard). -
test plan put <test-id>— replace a frontend test's plan-steps with a refined plan. Optional--expected-step-countdrift guard. -
project create/project update— manage projects from the CLI. Both commands pre-flight--target-urlagainst local addresses for fast feedback. -
test steps <test-id>— list a test's run steps with screenshot and DOM-snapshot pointers.--run-id <id>filters to the steps of one specific run. Without the flag, returns the cumulative step log across all runs with an advisory when steps span multiple runs. -
test result <test-id>— latest result: status, started/finished timestamps, video URL, step summary counts (passed / failed / skipped), and correlation fields (snapshotId,runId,codeVersion).--include-analysisadds an inline root-cause hypothesis, recommended fix target, and failure kind. -
test result <test-id> --history— list a test's prior runs (newest-first). Filters:--source cli|portal|mcp|schedule|github_action,--since 24h|7d|ISO,--page-size,--cursor. Each row carriesrunId,status,source,isRerun, timestamps,codeVersion, andfailureKind. A note is shown in place of a blank table for tests created before run-history tracking began. -
test failure get <test-id>— the agent entry point. Returns one self-consistent failure bundle: the failing step and its immediate neighbors with screenshots and DOM snapshots, the test source, the video pointer, a root-cause hypothesis, a recommended fix target, and correlation metadata. Every artifact in the bundle shares a singlesnapshotId; the CLI refuses to stitch data from different runs or code versions.--out <dir>writes the bundle atomically to disk.--failed-onlykeeps only the failing step and its neighbors. -
test failure summary <test-id>— one-screen triage card (status, failure kind, root-cause hypothesis, recommended fix target) without downloading media. -
test run <test-id>— trigger a fresh run. Without--wait, prints{ runId, status: "queued", … }and exits 0. With--wait, polls until terminal; exit 0 onpassed, exit 1 onfailed | blocked | cancelled, exit 7 on timeout with anextActionpointing attest wait <runId>. Accepts--target-url,--timeout,--idempotency-key. -
test run --all --project <id>— wave-ordered fresh batch run for all (or filtered) backend tests in a project. Routes to a batch endpoint; response enumeratesaccepted[],conflicts[],deferred[],skippedFrontend[], andskippedIntegration[]so a machine consumer readingacceptedalone can't silently undercount.--waitpolls all dispatched run IDs concurrently. -
test rerun [test-id…]— cheap replay of one or more tests. Frontend reruns replay the saved script verbatim (AI heal-on-drift is on by default, opt out with--no-auto-heal). Backend reruns expand the producer/teardown dependency closure; use--skip-dependenciesfor just the named test.--all --project <id>reruns every test in a project. Returnsaccepted[]plusdeferred[]for any tests shed by the per-key run-rate limit; under--wait, a non-emptydeferred[]exits 7 with a retry hint. -
test wait <run-id>— block until a run reaches a terminal status. Resumes polling after a timed-outtest run --wait, or when an agent already has arunId. Uses server-driven long-poll where supported; exponential backoff withRetry-Afterotherwise. -
test artifact get <run-id>— download the failure bundle for a specific run, addressed byrunIdinstead oftestId. Enforcesmeta.runId === <run-id>as an integrity check; exits 5 on mismatch. Default output directory:./.testsprite/runs/<run-id>/. -
--dry-run(global) — every command runs end-to-end without touching the network, credentials, or the local filesystem; emits canned data matching the API contract. -
Global flags:
--profile,--output json|text,--endpoint-url,--request-timeout <seconds>,--verbose,--debug. -
Pagination flags on every list:
--page-size,--starting-token,--max-items. -
--debugHTTP tracing to stderr (method, URL, request-id, latency, retry decisions). The API key is never included. -
Dashboard URL in outputs — commands that know both
projectIdandtestIdinclude adashboardUrldeep-link to the TestSprite web portal in JSON output. Text mode: create paths print aDashboard:line on stderr; run-completion output (test run --wait,test wait,test rerun --wait,test run --all) ends the run card with adashboardline on stdout. The portal domain is resolved from the configured API endpoint per environment. -
TTY-gated progress ticker — single-line in-place
\x1b[2K\rupdates during polling on TTY; completely silent on non-TTY (CI) and when--output jsonis set. -
AWS-CLI-style exit-code taxonomy — see Exit codes in README.
-
blockedis a distinct top-level status alongsidefailed(was collapsed intofailedin earlier previews). Triage routes:blocked→ infra (stale seed, login failure, unreachable target), not bug. -
test failure get/test stepsnow synthesize a terminalassertionstep row when no individual step is in error but the test failed at the assertion or overall-outcome layer. Previously the bundle shippedsteps: []for these tests. Synthetic rows have no screenshot or DOM snapshot. -
outcomeContributesToFailureboolean on every step row (nullwhen unclassified). The text renderer prefixes contributing rows with*so a 50-step list highlights which rows the failure landed on. -
failureKindenum widened: addsassertion_blocked,routing_404, andnetwork_timeout(previously these collapsed intounknown). The CLI accepts unrecognized values from the wire asunknownso new enum values are non-breaking. -
recommendedFixTargetreturnsnull(not anunknownwrapper) when the analysis pipeline produced no fill. Applied uniformly to/result?includeAnalysis,/failure, and/failure summary. -
Test.detailsdebug block ships a structuredprocessingStatus/testStatuspair alongside the previousrawStatusstring (deprecated but preserved for the transition window). -
test create --run/--wait/--timeout/--target-urlfully wired: chainsPOST /tests→POST /tests/{testId}/runs→GET /runs/{runId}in one invocation.--target-urlpre-flighted against local addresses on the client (exit 5) before the request is sent. -
Auto-minted idempotency keys and request IDs are suppressed by default; exposed under
--verbosefor retry and support use. -
Per-request wall-clock timeout (
--request-timeout, default 120 s) applied to every outgoing fetch. Under--wait, the per-request timeout is auto-raised to cover--timeout+ 5 s so a large batch under load is never cut at the default. -
test run --waitauto-resumes on 409run_in_flightby polling the existing run instead of exiting 6. An advisory is printed to stderr. Other conflict reasons and body-mismatch conflicts still propagate as exit 6. -
Backend test
test run --wait/test rerun --waitinclude a fallback path that readsGET /tests/{id}/resultwhen the run row is not yet finalized server-side, so the verdict is reachable without waiting for a timeout. -
test rerunbatch--waitsummary enumeratesdeferredandconflictscounts alongsidetotal(dispatched) so machine readers can't silently undercount.
-
parseEnvelopeBodynow recognizes NestJS raw 404 shape, surfacing the originalCannot POST /api/cli/v1/…message so the user sees which endpoint isn't deployed on the current backend rather than a generic "Server error." message. -
test run --waitCONFLICT auto-resume is gated ondetails.reason === 'run_in_flight'only. When--target-urlis supplied and the in-flight run's URL differs, the CLI fetches the existing run's URL and reports a descriptive conflict (exit 6) withnextAction: testsprite test wait <runId>. -
test stepsnow surfaces the synthetic terminalassertionstep row for assertion-only failures (previously this row was wired only fortest failure getandtest failure summary). -
test create-batch --plan-from-dir: theMAX_BATCH_SPECS(50) cap is enforced on valid specs after non-plan JSON files are skipped, not on the raw directory entry count. The duplicate-name advisory lookup uses a bounded 5-second deadline so a stalled listing endpoint delaystest createby at most 5 s. -
localValidationErrorandApiError.getDetail<T>()are shared library helpers; redundant inline cast patterns removed from call sites. -
engine-strict=truein.npmrcsonpm installhard-fails on Node < 20 instead of warning and proceeding. -
Commander
help [command]exits 0 (previously exited 5 ontest help/project help).