Skip to content

Commit 48b6eaf

Browse files
committed
feat: replace smoke baseline with curated 60-query Phase 1 dataset and integrate source-verification metadata in evaluation pipeline.
1 parent 2ae419e commit 48b6eaf

13 files changed

Lines changed: 4131 additions & 215 deletions

.DS_Store

2 KB
Binary file not shown.

README.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -478,7 +478,7 @@ See [reference_architecture.md](file:///Users/deepg/Desktop/KnowCode/docs/archit
478478
- ✅ Fix `metadata` type restriction (`dict[str, str]``dict[str, Any]`)
479479
- ✅ Harden configuration loading (logging, validation, strict server mode)
480480
- ✅ Decompose `KnowCodeService` and introduce `Protocol` interfaces
481-
- ✅ Add layer contract tests and harden retrieval evals (parser, store roundtrip, golden-query smoke baseline - see [docs/retrieval-evals.md](docs/retrieval-evals.md))
481+
- ✅ Add layer contract tests and harden retrieval evals (parser, store roundtrip, 60-query agent-curated baseline pending human review - see [docs/retrieval-evals.md](docs/retrieval-evals.md))
482482

483483
**Future releases:**
484484
- v2.4: Multi-level documentation synthesis (in progress: architecture/module/function export + freshness manifest)

docs/index.md

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -366,7 +366,7 @@ See [reference_architecture.md](file:///Users/deepg/Desktop/KnowCode/docs/archit
366366
- ✅ Fix `metadata` type restriction (`dict[str, str]``dict[str, Any]`)
367367
- ✅ Harden configuration loading (logging, validation, strict server mode)
368368
- ✅ Decompose `KnowCodeService` and introduce `Protocol` interfaces
369-
- ✅ Add layer contract tests and harden retrieval evals (parser, store roundtrip, golden-query smoke baseline — see [docs/retrieval-evals.md](retrieval-evals.md))
369+
- ✅ Add layer contract tests and harden retrieval evals (parser, store roundtrip, 60-query agent-curated baseline pending human review — see [docs/retrieval-evals.md](retrieval-evals.md))
370370

371371
**Future releases:**
372372
- v2.4: Multi-level documentation synthesis (in progress: architecture/module/function export + freshness manifest)
Lines changed: 173 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,173 @@
1+
# KnowCode — Re-Architecture Roadmap
2+
3+
*Consolidated from design discussion. Scope: scaling code-intelligence retrieval to very large (50M+ LOC) repositories under a tight memory budget while preserving strict correctness guarantees.*
4+
5+
---
6+
7+
## 0. Operating Constraints & Core Principle
8+
9+
**Target & budget**
10+
11+
- **Scale:** 50M+ LOC (~3–10M indexed chunks at symbol granularity).
12+
- **Memory:** ≤ 16 GB RAM for the query-serving process — *the binding constraint*.
13+
- **Latency:** p99 < 10 s; no per-query disk-I/O limit.
14+
- **Storage:** unconstrained.
15+
- **Freshness:** daily overnight incremental indexer; ≤ 24 h staleness acceptable.
16+
- **Authority:** the source tree is authoritative; correctness verified via robust invalidation, *not* per-query whole-tree comparison.
17+
18+
**The organizing principle everything hangs on: two contracts, one resolver.**
19+
20+
- **Exact layer (lossless, mechanically verifiable):** exact text + location come from reading the *live source tree* at query time. You cannot return stale text if you read it fresh.
21+
- **Empirical layer (best-effort discovery):** semantic/vector + lexical retrieval produce *candidate IDs only*. Every candidate is materialized through the exact layer before it is returned.
22+
23+
**Consequence:** the discovery layer can be arbitrarily lossy (quantization, stale embeddings) and still never compromise text correctness — it can only affect *which* chunks surface, never *what* a surfaced chunk says.
24+
25+
**Hard boundary established in discussion:** *"semantic search finds all and only relevant chunks 100% of the time" is not achievable.* Relevance is not mechanically defined, embeddings are lossy by construction, and turning a ranked list into a set requires an imperfect threshold. 100% precision+recall is reachable only when "relevant" becomes a **decidable predicate** (exact symbol / structural / regex queries) — which belongs in the *exact* layer, not the semantic one.
26+
27+
---
28+
29+
## 1. Current Solution (As-Built)
30+
31+
### 1.1 Embedding & Retrieval
32+
33+
- **Embedding model:** `voyage-code-3` (1024-dim, code-specialized), with correct asymmetric `input_type="document"` vs `"query"`. Provider cascade: explicit config → app config w/ API key → `DummyEmbeddingProvider` fallback.
34+
- **Dense arm:** FAISS `IndexFlatIP` — flat inner-product, i.e. brute-force, exact, **all vectors resident in RAM**.
35+
- **Sparse arm:** `InMemoryChunkRepository.search_by_tokens` — raw set-intersection overlap (`len(query_set & set(chunk.tokens))`). Labelled BM25 in docstrings; it is not.
36+
- **Fusion:** `HybridIndex.search` — Reciprocal Rank Fusion (RRF, K=60) with an `alpha=0.5` blend.
37+
- **Tokenizer:** code-aware — splits camelCase/snake_case, strips punctuation, lowercases, filters single-char tokens. Each chunk stores pre-computed `tokens: list[str]`.
38+
39+
### 1.2 Storage & Persistence
40+
41+
Entirely in-memory (plain Python dicts), serialized to flat files. **No database anywhere** in `src/`.
42+
43+
| File | Format | Contents |
44+
|---|---|---|
45+
| `chunks.json` | JSON | Chunk metadata: id, entity_id, content, tokens, metadata (embeddings excluded) |
46+
| `vectors.index` | FAISS binary | Dense flat-IP index |
47+
| `vectors.json` | JSON | id_map (int → chunk_id) + dimension |
48+
| `index_manifest.json` | JSON | Schema version, embedding & chunking config |
49+
| `knowcode_knowledge.json` | JSON | Entities + relationships graph |
50+
51+
On load, the **entire** state is deserialized and hydrated into memory before the first query.
52+
53+
### 1.3 What's Already Right (preserve)
54+
55+
- Code-specialized embedder with correct query/document asymmetry — getting this wrong silently craters recall.
56+
- Dense / sparse / RRF hybrid **architecture** already wired — the expensive structural work is done.
57+
- Metadata and vectors already split across separate files — **the migration seams are pre-drawn**.
58+
- Code-aware tokenization foundation.
59+
60+
---
61+
62+
## 2. Gaps Identified
63+
64+
| # | Gap | Category | Consequence |
65+
|---|---|---|---|
66+
| **G1** | FAISS flat = all vectors resident, brute-force | Memory/Scale | ~20 GB at 5M × 1024-dim float32 — blows the 16 GB budget |
67+
| **G2** | Full-JSON deserialization at startup | Startup | O(n) cold-start; everything hydrated before first query |
68+
| **G3** | No indexed token search at rest | Retrieval/Perf | Sparse arm linearly scans all chunks on every query |
69+
| **G4** | Sparse arm is set-intersection, not BM25 | Retrieval/Quality | No IDF / TF-saturation / length-norm; common tokens (`self`, `return`) drown rare identifiers (`reconcile_ledger`) — **destroys the exact-identifier signal that is the entire reason to run hybrid on code** |
70+
| **G5** | No lazy/partial loading | Scale | Cannot load a file/subtree without hydrating the universe |
71+
| **G6** | No concurrency/atomicity | Reliability | `json.dump` has no locking, journaling, or atomic-write semantics |
72+
| **G7** | Redundant source storage | Storage | Source text held in source tree *and* `chunks.json` (historically also knowledge.json) |
73+
| **G8** | Verbose JSON + absolute string IDs | Storage/Perf | Bloated artifacts, slow parse, no O(1) integer handles |
74+
| **G9** | Blend weight tuned against a broken sparse arm | Retrieval/Quality | `alpha=0.5` reflects compensating for noise, not real signal |
75+
| **G10** | No reranking stage | Retrieval/Quality | Fused top-k never refined by a deeper relevance signal |
76+
| **G11** | Location/freshness robustness not formalized | Correctness | Line-drift can mis-locate snippets between indexer runs if byte/line offsets are trusted blindly |
77+
78+
*G9–G11 are latent — they bite after the bigger items are fixed, or at scale.*
79+
80+
---
81+
82+
## 3. Candidate Solution Approaches
83+
84+
### 3.1 Metadata & Sparse Search → SQLite + FTS5
85+
86+
Replace the flat-file + in-memory-dict design for metadata with a single SQLite DB.
87+
88+
- **Real BM25** via FTS5's built-in `bm25()` (fixes G3, G4). IDF and length-normalization are present at FTS5 defaults — this alone kills all three of G4's sub-problems.
89+
- *Caveat:* FTS5 hardcodes k1=1.2, b=0.75 (not runtime-tunable). Fine for the bulk win; revisit only if measurement demands — then Tantivy (tunable BM25, separate index) or custom scoring over the `fts5vocab` table.
90+
- **Lazy row-level access** eliminates O(n) hydration (fixes G2, G5).
91+
- **WAL mode** + transaction-per-batch → atomic, crash-safe writes with concurrent reads (fixes G6) — exactly the "one overnight writer, many daytime readers" shape.
92+
- **Dense integer IDs** + JSON metadata column replace verbose absolute-ID JSON (fixes G8); the FAISS id_map folds into a `faiss_idx` column (one fewer file).
93+
- **Keep the code-aware tokenizer in Python:** pre-tokenize (camelCase/snake split, **plus index both subtokens *and* the original compound**`getUserById``get user by id getuserbyid`), space-join, feed to an **external-content** FTS5 table (`tokenize='unicode61'`). Your tokenization drives the BM25 stats; FTS5 does the inverted index + math; external-content = zero text duplication.
94+
- **Do NOT** store the FAISS index in SQLite as a BLOB — that reintroduces full-load. Vectors stay a separate mmap-friendly file.
95+
96+
### 3.2 Dense Vector Index → DiskANN/Vamana + RaBitQ
97+
98+
Replace FAISS flat (fixes G1).
99+
100+
- **Graph-based disk-resident index (Vamana / DiskANN family)** — current SOTA for large-scale, memory-frugal ANN: compressed vectors + graph in RAM, full-precision vectors on SSD; ~5–10× more points/node than HNSW at high recall. Matches the ≤16 GB-RAM / disk-unconstrained shape exactly.
101+
- **RaBitQ quantization** for the in-RAM tier — supersedes PQ: no codebook training, unbiased distance estimator with a rigorous error bound (PQ has none and fails on some datasets). At ~1 bit/dim, 1024-dim vectors → **~128 B each → ~0.64 GB resident at 5M chunks** (vs ~20 GB raw float32). Multi-bit variant trades memory for recall.
102+
- **Dynamic variants** (FreshDiskANN, LSM-VEC) matter: static graph indexes lose recall under repeated insert/delete — relevant to daily incremental updates. (OOD-DiskANN is notable — NL-query-vs-code-corpus is an out-of-distribution scenario.)
103+
- **Engines:** LanceDB (columnar + on-disk vectors, ships RaBitQ; aligns with the SQLite-style metadata design) or Qdrant (turnkey hybrid + rerank).
104+
105+
### 3.3 Retrieval Paradigm → Hybrid + Reranking
106+
107+
- **Hybrid (dense + real-BM25 sparse) via RRF** is SOTA consensus — beats single-method by ~5–15% generally, far more on recall in some settings. For code the win is outsized because BM25 nails exact identifiers that dense embeddings miss.
108+
- *Why nothing is currently on fire:* RRF is rank-only, so a bad arm's contribution is **bounded** to `1/(K+rank)` — which is why the broken sparse arm degrades gracefully and appears "carried by dense" rather than tanking the system.
109+
- **Re-tune the blend weight after fixing BM25** (fixes G9) — the optimum shifts toward sparse once it carries real signal. Same applies to k1/b if you move off FTS5: tune on code, not prose defaults.
110+
- **Reranking stage** on the fused top-k (fixes G10): cross-encoder (e.g. Voyage rerank-2.5) or late-interaction/ColBERT. The p99 < 10 s budget makes "retrieve top-1000 → rerank top-100/200" essentially free. *(Caveat: rerankers don't always help — measure.)*
111+
112+
### 3.4 Correctness → Source-Tree-Authoritative + Invalidation
113+
114+
- **Read exact text/location live from the source tree** at query time (fixes G7; this is the exact layer).
115+
- **Tiered location resolution** (addresses G11):
116+
1. *Fast path* — file content-hash matches indexed hash → trust stored byte offsets.
117+
2. *Slow path* — file changed → re-resolve symbol by qualified name.
118+
3. *Fallback* — content-fingerprint search; if the chunk is gone, report "changed/not found" rather than returning unverified text.
119+
- Use byte offsets over line numbers. **Invariant: never return text you can't mechanically verify is current.**
120+
- The token stream and pointers stored in SQLite are *derived index data*, not authoritative source — consistent with eventually dropping redundant `content`.
121+
122+
### 3.5 Exact / Decidable Query Engine (for queries that need 100% guarantees)
123+
124+
For the class of queries where "all and only" is required *and* decidable — exact symbol lookup, all-callers/callees, implementors of X, type-based queries, regex/literal search:
125+
126+
- **Structured indexes** (symbol table, call edges, type signatures, references) answer these with provable completeness — *conditional on the soundness of the per-language fact extractor*. Graph completeness is not assumed unless proven for the language/feature.
127+
- **Trigram index (Zoekt / Google Code Search pattern)** for literal/regex: trigrams yield a candidate superset with provably zero false negatives, then the regex runs against live source to remove false positives → **mechanically-provable 100% precision *and* recall**. Same philosophy as source-tree-authoritative: index narrows, source verifies. The trigram/lexical index also doubles as the sparse arm of hybrid.
128+
129+
### 3.6 Incremental Indexing
130+
131+
- **git diff** for changed-file detection (free, exact). **Content-hash** per file/chunk serves as *both* the invalidation key *and* the embedding-cache key — only genuinely changed chunks are re-embedded.
132+
- **Base + delta** index pattern: large base rebuilt periodically, small delta updated nightly, queries merge both, periodic compaction — bounds nightly cost.
133+
134+
---
135+
136+
## 4. Recommendations & Prioritization
137+
138+
Prioritized by **leverage ÷ effort**, sequenced so each ships independently (no big-bang rewrite). The existing three-file split makes these genuinely decoupled migrations.
139+
140+
| Priority | Action | Impact | Effort | Closes |
141+
|---|---|---|---|---|
142+
| **P0** | `SqliteChunkRepository` (same protocol) + FTS5 BM25 + WAL + lazy access | ★★★ | Low | G2, G3, G4, G5, G6, G8 |
143+
| **P1** | Re-tune RRF blend + add reranker (Voyage rerank-2.5) on fused top-k | ★★★ | Low | G9, G10 |
144+
| **P2** | `KnowledgeStore` → SQLite (nodes + edges, recursive CTEs) | ★★ | Med | graph load + subtree (rest of G2/G5) |
145+
| **P3** | FAISS flat → DiskANN/Vamana + RaBitQ (dynamic variant) | ★★★ | High | G1 |
146+
| **P4** | Source-tree-authoritative reads + tiered location resolution | ★★ | Med | G7, G11 |
147+
| **P5** | Decidable structured-query + trigram index (exact layer) | ★★ | High | enables 100%-guarantee queries |
148+
| **X-cut** | Incremental indexer: git-diff + content-hash embedding cache + base/delta | ★★ | Med | freshness / re-embed cost |
149+
| **X-cut** | Measurement: ablate dense-only vs hybrid vs BM25-hybrid on held-out real queries (nDCG@10, recall@50) || Low | validates every change above |
150+
151+
**Why P0 is first.** Single highest-leverage, lowest-effort move: converts the placeholder sparse arm into real BM25 (your fastest quality jump) *and* retires the startup, lazy-load, and concurrency gaps in one migration. It also unblocks the trigram/exact-layer work (P5), since the same DB hosts both.
152+
153+
**Why reranking is P1, not later.** Cheap given the latency budget, stacks directly on the existing RRF output, and combined with P0 likely captures most of the achievable discovery-quality gain *before* the heavier vector migration.
154+
155+
**Why the vector migration (P3) is high-effort/late despite fixing the headline memory gap.** It is the largest change and is *orthogonal* to everything else (separate file, separate concern). Doing it after P0–P1 means you are not debugging two subsystems at once, and you will have a measurement harness in place to confirm recall is preserved post-quantization.
156+
157+
**Cross-cutting discipline.** Validate every change empirically on a held-out set of real queries from your own repos — because code-retrieval benchmarks (CoIR et al.) are themselves noisy (label noise, tasks that reduce to string matching), the only trustworthy signal is your own corpus.
158+
159+
---
160+
161+
## 5. Key Techniques & References
162+
163+
- **RaBitQ** — Gao & Long, *Quantizing High-Dimensional Vectors with a Theoretical Error Bound for ANN Search* (arXiv:2405.12497); multi-bit extension (2025), GPU version (2026), RaBitQ Library. PQ-superseding quantization with error bounds.
164+
- **DiskANN / Vamana** — Subramanya et al., *Fast Accurate Billion-point NN Search on a Single Node* (NeurIPS 2019). Disk-resident graph ANN. Dynamic: FreshDiskANN; LSM-VEC (arXiv:2505.17152); OOD-DiskANN.
165+
- **FTS5 / BM25** — SQLite FTS5 full-text extension with built-in `bm25()` ranking (`sqlite.org/fts5.html`).
166+
- **Hybrid retrieval / RRF** — sparse+dense fusion via Reciprocal Rank Fusion; SOTA consensus that hybrid beats single-method, especially for identifier-heavy code.
167+
- **Late interaction / ColBERT** — token-level reranking; strong precision. Cross-encoder rerankers (e.g. Voyage rerank-2.5).
168+
- **Code embedding & benchmarks**`voyage-code-3` (VoyageAI); CoIR benchmark (`github.com/CoIR-team/coir`); CodeXEmbed (arXiv:2411.12644); Qodo-Embed.
169+
- **Trigram code search** — Zoekt / Google Code Search; provably-complete regex candidate generation, verified against live source.
170+
171+
---
172+
173+
*Sequencing in one line: **P0 → P1** first (real BM25 + rerank — biggest quality-per-effort), then **P2** (graph → SQLite), then **P3** (DiskANN/RaBitQ for the memory ceiling), then **P4–P5** (formal correctness + decidable exact-query layer). Validate each step on your own held-out queries.*

0 commit comments

Comments
 (0)