Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 12 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -71,6 +71,18 @@ python app.py --base-dir /path/to/claude/projects
> The server **refuses to start** if `--debug` is combined with a non-loopback `--host` (e.g. `0.0.0.0`).
> That check runs only when you start the app with **`python app.py`** (not via `flask run` or other WSGI entrypoints).

#### Deployment — WSGI workers
Comment thread
clean6378-max-it marked this conversation as resolved.
Outdated

The FTS search index (`utils/search_index.py`) assumes **a single Python process** serves the app:

- One daemon thread rebuilds the index and publishes it by atomically swapping `search_index.active` (see [Search index — locks and threading](docs/architecture.md#search-index-fts5--locks-and-threading)).
- Module-level locks and the in-memory usability cache are **per process**, not shared across workers.
- A cross-process advisory lock (`search_index.background.lock`) ensures at most **one** background refresher per index directory on a host, but **does not** coordinate search reads or caches across workers.

**Do not run multiple Gunicorn/uWSGI workers** (`gunicorn -w 2`, etc.) against one operator-facing deployment. Requests will hit different processes with divergent cache state; only one worker runs background refresh; index freshness and `index_locked` / live-scan fallback become nondeterministic. Use **`--workers 1`** (or run `python app.py`) for supported behavior. For horizontal scale, run separate single-worker instances with disjoint `CLAUDE_CODE_CHAT_BROWSER_SEARCH_INDEX_DIR` and workspace paths — not multiple workers behind one load balancer on the same paths.

Disable the index entirely with `CLAUDE_CODE_CHAT_BROWSER_NO_SEARCH_INDEX=1` if you only need live JSONL scan (slower, but avoids index files).
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated

### CLI Export

```bash
Expand Down
52 changes: 51 additions & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@
3. User opens a project → `GET /api/projects/<name>/sessions` → full `parse_session()` per file, exclusion filter, summary rows.
4. User opens a session → `GET /api/sessions/<project>/<id>` → full session JSON for the message panel.
5. Optional: `GET /api/sessions/.../stats` for sidebar metrics without loading all messages.
6. Search: `GET /api/search?q=...` scans all projects (brute force).
6. Search: `GET /api/search?q=...` queries the local FTS index when usable, with live JSONL scan fallback (see [Search index](#search-index-fts5--locks-and-threading)).
7. Export: `POST /api/export` or `GET /api/export/session/...` → Markdown/zip via exporters; state file updated on successful bulk export.

## Dispatch table
Expand Down Expand Up @@ -143,6 +143,56 @@ No bundler step — modern browsers load modules directly. Frontend unit tests u
- `integration-tests` — API integration subset + coverage artifact
- `js-tests` — `npm ci` + vitest

## Search index (FTS5) — locks and threading
Comment thread
clean6378-max-it marked this conversation as resolved.

`utils/search_index.py` maintains a local SQLite FTS5 index under `~/.claude-code-chat-browser/` (override with `CLAUDE_CODE_CHAT_BROWSER_SEARCH_INDEX_DIR`). Session JSONL on disk remains the source of truth; `GET /api/search` can fall back to a live scan when the index is missing, stale, or temporarily locked during rebuild.

### Threading model

| Role | Who | What |
|------|-----|------|
| **Readers** | Flask request threads (and synchronous callers such as `ensure_search_index` on the first request path) | Open the **active** index DB read-only via `_resolve_active_index_db_path()`; `query_index_hits` does not take any module `threading.Lock`. |
Comment thread
clean6378-max-it marked this conversation as resolved.
Outdated
| **Writer** | One daemon thread (`search-index-refresh`, started from `start_search_index_background` in `app.py`) | Periodically calls `ensure_search_index` → `build_search_index` when the projects fingerprint changes. |
| **Publish** | Writer, at end of a successful rebuild | `_publish_active_index` writes a new `search_index.<id>.sqlite`, atomically replaces `search_index.active` (write temp + `Path.replace`), then prunes older DB files. Readers always open whichever DB name the pointer names — they see the **previous** or **new** index, never a half-written DB file. |
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated

Rebuild work happens on a **new** SQLite file; the active pointer switches only after the new file is committed. Concurrent FTS reads against the old file may fail with `sqlite3.OperationalError` (for example database locked); `query_index_hits` surfaces that as `index_locked=True` so `/api/search` can degrade to live scan instead of returning a silent empty success.
Comment thread
coderabbitai[bot] marked this conversation as resolved.
Outdated

### Lock inventory

Four module-level synchronization primitives guard index lifecycle and caches (`utils/search_index.py`):

| Primitive | Type | Guards |
|-----------|------|--------|
| `_index_lock` | `threading.Lock` | `_background_started` and startup/teardown of the background worker (`start_search_index_background`, `reset_background_for_tests`). |
| `_index_build_lock` | `threading.Lock` | Entire `build_search_index` body — at most one rebuild at a time **in this process**. |
| `_background_lock_fd` | OS advisory lock on `search_index.background.lock` | Cross-process **single background owner** for a shared index directory (`_try_acquire_cross_process_background_lock`). |
| `_usability_cache_lock` | `threading.Lock` | In-memory `_usability_cache` used by `index_is_usable` and cleared after rebuild (`_clear_usability_cache`). |

**Usage map (current code):**

- `_index_build_lock` — wraps `build_search_index` (serializes rebuild and fingerprint short-circuit reads).
- `_index_lock` — held briefly while deciding whether to start the daemon and while resetting test state; may call `_try_acquire_cross_process_background_lock` while held.
- `_background_lock_fd` — non-blocking exclusive flock/msvcrt lock; failure means another process already owns background refresh for this cache dir.
- `_usability_cache_lock` — `index_is_usable` and `_clear_usability_cache`.
- `_publish_active_index` — pointer swap + `_prune_stale_index_files` (no module lock; relies on atomic pointer replace and build serialization).

### Lock-acquisition contract (must stay so)

The four primitives are **not nested arbitrarily**. New code must preserve this contract; the planned Week 30 concurrency suite (`tests/test_search_index_concurrency.py`, PR after this doc) will assert it and fail if lock order regresses.

1. **`_index_lock`** may be taken alone, or together with acquiring the advisory background lock during `start_search_index_background`. Do **not** call `build_search_index` while holding `_index_lock`.
2. **`_index_build_lock`** is taken only inside `build_search_index`. While held, code may open the active DB read-only and write a **new** sidecar SQLite file. The only permitted nested module lock is **`_usability_cache_lock`** via `_clear_usability_cache` at the end of a successful rebuild (`_index_build_lock` → `_usability_cache_lock`).
3. **`_usability_cache_lock`** must never be held while waiting on `_index_build_lock` or `_index_lock`. Typical pattern: short cache read/update in `index_is_usable` without holding the build lock.
4. **`_background_lock_fd`** is independent of the three `threading.Lock` instances except at startup: acquire the advisory lock only from `_try_acquire_cross_process_background_lock`, which is called under `_index_lock` once per process.

**Forbidden:** holding `_usability_cache_lock` and then acquiring `_index_build_lock` or `_index_lock`; holding `_index_lock` across `build_search_index`; taking `_index_build_lock` from multiple threads without going through the existing `build_search_index` entry point.

### Deployment assumptions

The design targets **one WSGI worker process** (for example `python app.py` or `gunicorn --workers 1`). That gives one in-process writer thread, one advisory-lock owner, and consistent in-memory usability cache per operator session.

See [Deployment — WSGI workers](../README.md#deployment-wsgi-workers) in the README for multi-worker warnings.
Comment thread
clean6378-max-it marked this conversation as resolved.
Outdated

## What this codebase is not

- **Not multi-user** — no authn/authz; single local operator.
Expand Down
Loading