Skip to content

Commit 612578a

Browse files
committed
docs: add SDLC document-collateral retrieval strategy
Design proposal extending KnowCode retrieval from code to SDLC prose collateral (PRDs, HLD/LLD, release notes, case studies, whitepapers). Captures the four interview decisions (hybrid deployment, LLM-RAG consumer, all query patterns, templated docs), a D1-D9 gap register, a P0-P5 leverage/effort ladder, the document/traceability graph schema, a prose eval plan, and a reference-architecture layer mapping. Scale is a non-issue (~500K chunks); budget goes to quality.
1 parent ed69efe commit 612578a

1 file changed

Lines changed: 252 additions & 0 deletions

File tree

Lines changed: 252 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,252 @@
1+
# KnowCode — SDLC Document-Collateral Retrieval Strategy
2+
3+
*Consolidated from design discussion (June 2026). Scope: extending KnowCode's retrieval substrate from **code** to **SDLC prose collateral** — Problem Statements, User Personas, PRDs, Architecture/HLD/LLD, Release Notes, User/Installation Guides, Support Manuals, Marketing Brochures, Product Feature Catalogs, Client Case Studies, Whitepapers — to maximize semantic-search and RAG quality over thousands of templated markdown documents, each thousands of lines long.*
4+
5+
> **Status:** Design proposal — not yet implemented
6+
> **Owner:** Solo
7+
> **Companion to:** [`reference_architecture.md`](../architecture/reference_architecture.md) · [`knowcode-architecture-synthesis.md`](./knowcode-architecture-synthesis.md) · [`retrieval-evals.md`](../retrieval-evals.md)
8+
> **`[ASPIRATIONAL]`** items are target design, not shipped features — consistent with the `[HARDENED]` convention in the reference architecture.
9+
10+
---
11+
12+
## 0. Operating Constraints & Core Principle
13+
14+
**The four decisions that shaped this design** (from the requirements interview):
15+
16+
| Decision | Choice | Consequence |
17+
|---|---|---|
18+
| Deployment / privacy | **Hybrid** | Local storage + index; hosted embedding/rerank APIs for non-sensitive collateral; **local models for confidential docs** (PRDs, LLDs). LLM-in-the-loop indexing is permitted. |
19+
| Primary consumer | **LLM / agent RAG** | Optimize **recall + context sufficiency + token budget**, not human-facing precision@k UX. Feeds the existing sufficiency-score → local-vs-frontier gate. |
20+
| Query patterns | **All three** — fact lookup, cross-doc traceability, synthesis/comparison | Forces an **adaptive query router** over a shared substrate, not a single retrieval path. |
21+
| Corpus structure | **Highly templated + rich frontmatter** | Structure is exploitable: section-aware chunking, metadata-filtered hybrid, mechanical cross-doc linking by `feature_id`/`requirement_id`. |
22+
23+
**The organizing principle: scale is not the constraint — quality is.**
24+
25+
- ~5,000 docs × ~3,000 lines ≈ **100–200M tokens ≈ ~500K chunks**. A flat FP32 index at 1024-dim is ~2 GB resident. It fits in RAM on a developer laptop.
26+
- **Therefore the entire DiskANN/Vamana + RaBitQ machinery from [`knowcode-architecture-synthesis.md`](./knowcode-architecture-synthesis.md) is out of scope here.** Keep an exact flat index (zero quantization recall loss). Spend the whole complexity budget on retrieval quality: chunking, embeddings, context preservation, metadata, the document graph, and reranking.
27+
28+
**The prose analogue of "the source tree is authoritative": *document structure is the oracle of first resort.*** Because the corpus is templated, headings, frontmatter, and feature/requirement IDs are mechanically reliable signals — they are exploited *before* any embedding similarity is trusted, exactly as the code system trusts the AST before the vector arm.
29+
30+
---
31+
32+
## 1. Current Solution (As-Built, code-tuned)
33+
34+
What exists today is built for code and inherits the wrong defaults for prose.
35+
36+
### 1.1 Retrieval
37+
- **Embedding:** `voyage-code-3` (1024-dim, **code-specialized**), correct asymmetric `document`/`query` input types.
38+
- **Dense arm:** FAISS `IndexFlatIP` — exact, brute-force, RAM-resident.
39+
- **Sparse arm:** in-flight migration to **SQLite FTS5 BM25 + WAL** (P0 of the code re-arch; replaces the placeholder set-intersection arm).
40+
- **Fusion:** RRF (K=60, `alpha=0.5`) — tuned against the *old* sparse arm.
41+
- **Rerank:** VoyageAI cross-encoder with signal-based fallback.
42+
- **Expansion:** dependency-aware expansion over the NetworkX semantic graph (call/import/contains/inherits).
43+
44+
### 1.2 Eval (the discipline to preserve)
45+
- Golden set `golden_v1.0.json` (60 queries), three-role **Author / Oracle / Adversary** pipeline, mechanical validation gate.
46+
- Gates: Mean MRR ≥ 0.40, Recall@10 ≥ 0.50, File-Coverage@5 ≥ 0.50, easy-query P@1 = 1.0, LOCATE MRR ≥ 0.60, no zero-recall task type.
47+
- **Known finding:** `sufficiency_score` is over-confident on the code corpus — `answer_correctness@0.8 = 0.571` across 7 routed records. Any prose port must re-measure this, not assume it.
48+
49+
### 1.3 What carries over (≈70% of the investment)
50+
FAISS flat index · SQLite FTS5 BM25 · RRF (re-tuned) · VoyageAI rerank (upgraded) · NetworkX graph (extended) · task-type classifier (→ router) · three-role golden pipeline · sufficiency-score gate (re-calibrated).
51+
52+
---
53+
54+
## 2. Gaps for Prose (Gap Register)
55+
56+
| # | Gap | Category | Consequence |
57+
|---|---|---|---|
58+
| **D1** | `voyage-code-3` is code-specialized | Embedding/Quality | Wrong vector space for natural-language collateral; recall craters on prose intent |
59+
| **D2** | Symbol/AST chunking doesn't apply; fixed windows would shred templated structure | Chunking/Quality | Splits tables, mixes unrelated sections, destroys the structure signal the corpus hands you for free |
60+
| **D3** | Chunks lose document context | Context/Recall | A section of a 3,000-line PRD is ambiguous out of context; the dominant prose failure mode |
61+
| **D4** | No metadata layer (`doc_type`, `product`, `version`, `audience`, `status`, dates) | Filtering/Precision | Cannot filter, cannot answer comparison/synthesis ("release notes v2 vs v3", "all healthcare case studies") |
62+
| **D5** | Graph models code entities, not doc/feature/requirement entities | Graph/Capability | No cross-doc traceability — the highest-value query class is structurally unanswerable |
63+
| **D6** | RRF/`alpha` tuned for identifier-heavy code | Fusion/Quality | Prose has less exact-token signal; blend optimum shifts dense-heavier |
64+
| **D7** | No synthesis/abstraction layer | Synthesis/Capability | "Summarize" / "what changed" require map-reduce over many chunks; flat top-k can't |
65+
| **D8** | Reranker + sufficiency-score calibrated on code | Eval/Calibration | Unknown behavior on prose; the 0.8 over-confidence finding may differ in sign and size |
66+
| **D9** | No prose golden set | Eval/Process | No measurement, no regression gate, no calibration target for D1–D8 |
67+
68+
---
69+
70+
## 3. Candidate Solution Approaches
71+
72+
### 3.1 Structure-aware parent-child chunking *(fixes D2, D3)*
73+
- Chunk on the **templated heading hierarchy** (H1→H2→H3), never fixed token windows. Keep tables, fenced code, mermaid, and lists **atomic**.
74+
- **Small-to-big / parent-document retrieval:** embed the small leaf section (precision); return the parent section or whole doc (context).
75+
- **Per-doc-type section schema:** tag each chunk with its canonical role from the template (a PRD `Non-Goals``Requirements`) → enables section-targeted retrieval and is the join key for the graph (§3.5).
76+
- Reuse and extend the existing markdown parser (already emits structural entities).
77+
78+
### 3.2 Prose embedding tier *(fixes D1)*
79+
- **Cloud (non-sensitive):** `voyage-3-large` — top RAG tier, *same Voyage SDK already wired in* (minimal change). Alternatives: `Cohere embed-v4` (multimodal — diagrams/screenshots as images), `Gemini Embedding` (top MTEB).
80+
- **Local (confidential PRDs/LLDs):** `Qwen3-Embedding` (8B/4B) — SOTA open weights, Matryoshka truncation, on-machine.
81+
- One vector space per tier; tag chunk provenance with the embedding model. Do **not** mix embedders in one index.
82+
83+
### 3.3 Context preservation *(fixes D3 — the recall multiplier)*
84+
- **Contextual Retrieval (Anthropic):** prepend a 50–100-token LLM-generated blurb to each chunk before embedding ("This is the rate-limiting section of the v2 Atlas PRD, which specifies…"). ~35% fewer retrieval failures (49% paired with contextual BM25). **Prompt-cache the doc**, generate per-chunk context → ~$1/M doc-tokens; run context generation on the **local model** for sensitive docs.
85+
- **Cheaper zero-LLM alternatives:** `voyage-context-3` contextualized-chunk embeddings, or **late chunking** (embed the long doc, pool per-chunk). Less lift than full contextual retrieval; no per-chunk LLM call.
86+
87+
### 3.4 Metadata-filtered hybrid *(fixes D4, D6)*
88+
- Reuse the in-flight SQLite FTS5 store; add columns `doc_type, product, version, date, audience, status, feature_id`. **Filter-then-search** collapses the candidate set before ANN — large precision win and the backbone of every synthesis/comparison query.
89+
- Keep BM25 in the blend (prose still has exact tokens dense misses — product names, version strings, error codes, acronyms). **Re-tune RRF weights dense-heavier than the code config; measure, don't guess.**
90+
91+
### 3.5 Document / traceability graph *(fixes D5 — the moat)*
92+
Extend NetworkX; do not rebuild. See §5 for the schema. Auto-builds a **requirements-traceability matrix** so "trace Feature X" is a graph walk, not a vector search. 2026 research in this exact space: **SAGE** (structure-aware graph expansion over heterogeneous data), **Orion-RAG** (path-aligned hybrid).
93+
94+
### 3.6 Reranking, prose-grade *(fixes D8 partial)*
95+
- Retrieve top-100–200 → rerank → top-k; the p99 budget makes this near-free and it's the #1 precision lever after chunking.
96+
- Models (Feb 2026 ELO leaders): `Zerank-2`, `Cohere Rerank 4 Pro`; stay in-house with `Voyage rerank-2.5` (`-lite` halves latency). Local/sensitive: `bge-reranker-v2-m3`.
97+
- **Skip ColBERT** — 2026 consensus: bi-encoder + cross-encoder is simpler and quality-equivalent for this.
98+
99+
### 3.7 Synthesis & abstraction *(fixes D7)*
100+
- **RAPTOR:** recursively cluster+summarize chunks into a per-doc (and per-collection) tree; retrieve at the right altitude. Natural fit for thousand-line docs and the existing Layer-7 abstraction-levels ambition.
101+
- **LazyGraphRAG (Microsoft 2025):** community summaries + map-reduce over the graph, built at query time at ~0.1% of full-GraphRAG indexing cost. Fits the cost-aware hybrid posture.
102+
103+
### 3.8 Query router + gated agentic loop *(ties it together)*
104+
- Reuse the task classifier as a **query router** (adaptive routing — match query complexity to pipeline complexity). Keep the cheap path cheap: agentic retrieval is waste on simple factual queries.
105+
- Hard multi-hop: HyDE, multi-query/RAG-Fusion, decomposition; controlled-vocab expansion from the Feature Catalog/glossary. Wrap in **Self-RAG / Corrective-RAG gated by sufficiency score**: retrieve → grade → re-retrieve → escalate to frontier LLM only when score < threshold.
106+
107+
---
108+
109+
## 4. Recommendations & Prioritization
110+
111+
Prioritized by **leverage ÷ effort**, sequenced so each ships independently. The existing three-index seams (vector / FTS / graph) make these decoupled migrations, exactly as in the code re-arch.
112+
113+
| Priority | Action | Impact | Effort | Closes |
114+
|---|---|---|---|---|
115+
| **P0** | Structure-aware parent-child chunking + per-doc-type section schema | ★★★ | Low | D2, D3 |
116+
| **P1** | Prose embedding tier (`voyage-3-large` cloud / `Qwen3-Embedding` local) + contextual retrieval | ★★★ | Low–Med | D1, D3 |
117+
| **P2** | Metadata-filtered hybrid (reuse SQLite FTS5) + re-tuned RRF + prose reranker | ★★★ | Low | D4, D6, D8 |
118+
| **P3** | Document / traceability graph (extend NetworkX) | ★★★ | Med–High | D5 |
119+
| **P4** | Query router (reuse classifier) + adaptive paths | ★★ | Med | binds P0–P5 |
120+
| **P5** | Synthesis layer — RAPTOR + LazyGraphRAG | ★★ | High | D7 |
121+
| **X-cut** | **Prose golden set** + metrics + sufficiency re-calibration (§6) || Med | D8, D9 |
122+
| **X-cut** | Agentic/iterative retrieval gated by sufficiency || Med | hard multi-hop |
123+
124+
**Why P0 first.** Structure-aware chunking is the largest single prose-quality lever and unblocks everything downstream (the section schema is the graph's join key). **Why P1–P2 next.** They capture most achievable discovery quality before any graph work — the 2026 field's own top-five ordering is chunking → hybrid → rerank → query-transform → eval. **Why the graph (P3) is the differentiator despite higher effort.** It is the one capability flat RAG can never replicate and it serves all three query classes (entry-point for fact lookup, traversal for traceability, subgraph scoping for synthesis).
125+
126+
**If you do only five things:** P0 chunking · P1 prose embedder · P1 contextual retrieval · P2 metadata-hybrid + reranker · P3 the graph — then prove all five on the prose golden set before tuning.
127+
128+
---
129+
130+
## 5. The Document / Traceability Graph (schema)
131+
132+
Extend the existing entity/relationship model. New entity kinds and edges:
133+
134+
```yaml
135+
# New entity kinds (alongside code function|class|module|...)
136+
DocEntity:
137+
kinds: [document, section, feature, requirement, persona,
138+
component, release, decision, client]
139+
source_location: { file, heading_path, line_range }
140+
doc_type: problem_statement | prd | architecture | hld | lld |
141+
release_notes | user_guide | install_guide | support_manual |
142+
brochure | feature_catalog | case_study | whitepaper
143+
metadata: { product, version, date, audience, status, feature_id, requirement_id }
144+
embeddings: vector # contextualized chunk embedding
145+
content_hash: sha256 # rename/edit-resilient identity (reuse AD-7 pattern)
146+
147+
# New edges
148+
Edges:
149+
structural: document --contains--> section
150+
explicit: section --links_to--> {section|document} # markdown links, ticket IDs
151+
document --implements--> requirement # from frontmatter
152+
document --supersedes--> document # version lineage
153+
derived: {prd,hld,lld,release} --traces--> feature # mechanical via feature_id
154+
requirement --satisfied_by--> component # PRD→HLD join
155+
component --detailed_in--> section # HLD→LLD join
156+
feature --documented_in--> {guide,case_study} # downstream collateral
157+
```
158+
159+
**Traceability walk (the headline capability):**
160+
```
161+
Problem Statement → PRD requirement → HLD component → LLD module
162+
→ Release Note → User Guide → Support article
163+
```
164+
Because the corpus is templated, the `feature_id`/`requirement_id` joins are **mechanical**, not LLM-inferred — the same "structure is the oracle" guarantee that makes the code AST trustworthy. Entity-link doc mentions to canonical nodes in the **Feature Catalog / Persona list**; that linkage doubles as the controlled vocabulary for query expansion (§3.8).
165+
166+
---
167+
168+
## 6. Evaluation Plan (prose golden set)
169+
170+
Reuse the three-role Author / Oracle / Adversary pipeline verbatim; the **Oracle reads documents instead of code**, and the mechanical gate resolves `(doc_id, heading_path, line_range)` and `feature_id` joins instead of AST symbols.
171+
172+
**Stratification:** `doc_type × query_type × difficulty`, where `query_type ∈ {fact_lookup, traceability, synthesis}` (replacing the code task types). Traceability and synthesis strata must be present from Phase 1 — they are where flat RAG fails and where the graph earns its keep.
173+
174+
**Record schema additions** (extend the existing `GoldenLabel`):
175+
```json
176+
{
177+
"query_type": "traceability",
178+
"expected_doc_spans": [{"doc_id": "...", "heading_path": "PRD > Requirements > Rate Limiting", "line_range": [..]}],
179+
"expected_trace_path": ["feature:atlas-ratelimit", "prd:...", "hld:...", "release:..."],
180+
"must_mention_facts": ["..."],
181+
"must_not_mention_facts": ["..."]
182+
}
183+
```
184+
185+
**Metrics (RAG-consumer-tuned):**
186+
- Retrieval: **recall@k is primary** (a missed chunk = a wrong generated answer), plus nDCG@10, MRR.
187+
- Traceability: **path-completeness** — fraction of the expected trace path recovered.
188+
- Generation: **faithfulness / groundedness + citation accuracy** (RAGAS-style) over the assembled context.
189+
- **Re-calibrate `sufficiency_score`** against prose answer-correctness; publish the reliability diagram. Do not port the 0.8 threshold — re-measure (the code finding was *over*-confident; prose may differ in sign).
190+
191+
**Field target to beat:** advanced RAG ≈ **63% factual accuracy vs ≈ 44% naive** (2026). The ladder in §4 is *how*; this harness is *how you prove* each rung.
192+
193+
---
194+
195+
## 7. Mapping onto the Reference Architecture
196+
197+
This reuses the existing layered design — it is an extension, not a parallel system.
198+
199+
| Layer | Extension for prose collateral |
200+
|---|---|
201+
| L1 Ingestion | Frontmatter + doc-type detection; provenance per document/version |
202+
| L2 Parsing | Markdown heading-tree + table/code/diagram preservation (extend existing markdown parser) |
203+
| L3 Semantic Graph | New doc/feature/requirement entities + traceability edges (§5) |
204+
| **L4a Search/Index** | Prose embedder, parent-child chunks, contextual retrieval, metadata-filtered hybrid, re-tuned RRF, prose reranker |
205+
| L7 Doc Synthesis | RAPTOR trees feed multi-level summaries (Executive→…→Section) |
206+
| L9 Context Synthesis | Small-to-big assembly, per-chunk provenance headers (`[PRD v2 · §Rate Limiting · 2026-03]`), token budget |
207+
| L10a/b Agent + MCP | Query router; new tools — `trace_feature`, `compare_docs`, `summarize_collection`; sufficiency re-calibrated |
208+
| L11 Feedback | Prose golden set + drift detection (§6) |
209+
210+
```mermaid
211+
flowchart TB
212+
Q[Query] --> R{Query Router<br/>reuse task classifier}
213+
R -->|fact lookup| H[Hybrid: dense + BM25 + RRF<br/>+ metadata pre-filter]
214+
R -->|traceability| G[Graph walk<br/>feature/req IDs across SDLC chain]
215+
R -->|synthesis| T[RAPTOR trees +<br/>LazyGraphRAG summaries]
216+
H --> RR[Cross-encoder rerank]
217+
G --> RR
218+
T --> RR
219+
RR --> SB[Small-to-big parent expansion<br/>+ provenance headers]
220+
SB --> SC{Sufficiency score}
221+
SC -->|>= threshold| LOC[Local answer · 0 frontier tokens]
222+
SC -->|< threshold| LLM[Frontier LLM w/ grounded context]
223+
224+
subgraph Substrate[Shared substrate]
225+
VEC[(Flat FAISS · voyage-3-large<br/>contextual chunks)]
226+
FTS[(SQLite FTS5 · BM25 + metadata)]
227+
GR[(Doc/Feature graph · NetworkX)]
228+
end
229+
H -.-> VEC
230+
H -.-> FTS
231+
G -.-> GR
232+
T -.-> GR
233+
T -.-> VEC
234+
```
235+
236+
---
237+
238+
## 8. Key Techniques & References
239+
240+
- **Contextual Retrieval** — Anthropic (Sept 2024). Per-chunk LLM context prefix; ~35%/49% fewer retrieval failures; prompt-cache to amortize.
241+
- **Late Chunking** — Jina (2024). Long-context embed-then-pool; zero-LLM context preservation.
242+
- **Parent-document / small-to-big retrieval** — embed small, return big; standard 2026 production pattern.
243+
- **RAPTOR** — recursive clustering + abstractive summary tree; multi-altitude retrieval over long docs.
244+
- **GraphRAG / LazyGraphRAG** — Microsoft (2024 / 2025). Community summaries; Lazy variant at ~0.1% indexing cost.
245+
- **SAGE** (arXiv:2602.16964) — structure-aware graph expansion for heterogeneous data. **Orion-RAG** (arXiv:2601.04764) — path-aligned hybrid retrieval. Both directly relevant to the §5 graph.
246+
- **Embedders (2026):** `voyage-3-large`, `Cohere embed-v4` (multimodal), `Gemini Embedding` (MTEB-top), `Qwen3-Embedding` (SOTA open, MRL).
247+
- **Rerankers (2026):** `Zerank-2`, `Cohere Rerank 4 Pro`, `Voyage rerank-2.5`/`-lite`, `bge-reranker-v2-m3` (open).
248+
- **Hybrid + RRF, Self-RAG, Corrective-RAG (CRAG), HyDE, RAG-Fusion** — query-adaptive retrieval orchestration.
249+
250+
---
251+
252+
*Sequencing in one line: **P0 chunking → P1 prose embedder + contextual retrieval → P2 metadata-hybrid + reranker** (most quality-per-effort), then **P3 the traceability graph** (the differentiator), then **P4 router → P5 synthesis**. Stand up the prose golden set in parallel and validate every rung on it — scale was never the problem; measured quality is the whole game.*

0 commit comments

Comments
 (0)