Pinned Loading
-
-
ai-systems-notes/local-llm-rag-cag-benchmark
ai-systems-notes/local-llm-rag-cag-benchmark PublicローカルLLM(RTX 4070 + vLLM / Qwen3.5-4B-FP8)で RAG と CAG を同一条件で比較する、再現可能な日本語ベンチマーク。自作小説50問を正答率・TTFT・トークンコストで実測。
Python
-
ai-systems-notes/ollama-prefill-kv-restore
ai-systems-notes/ollama-prefill-kv-restore PublicOpt-in KV-cache (prefill) save/restore for Ollama — reproducible TTFT benchmark, up to 417× on a fixed system prompt. 固定 system prompt の再プレフィルを KV 復元で置き換え、TTFT を最大 417× 短縮。llama3.2:3b / RTX 4070、3回…
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.

