Skip to content

Commit 0e1b029

Browse files
docs: update README with multi-model benchmark data (1,300+ evals, 10 models, 3 providers)
1 parent 3ba2092 commit 0e1b029

1 file changed

Lines changed: 10 additions & 25 deletions

File tree

README.md

Lines changed: 10 additions & 25 deletions
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@
77

88
Python implementation of [GCF](https://gcformat.com/) — the most token-efficient wire format for LLMs. A drop-in alternative to JSON and TOON for any structured data.
99

10-
**79% fewer input tokens than JSON. 75% fewer output tokens. 52% smaller than TOON. 100% LLM comprehension at 500 symbols, where JSON scores 76.9% and TOON scores 92.3%.**
10+
**79% fewer input tokens than JSON. 63% fewer output tokens. 90.5% average comprehension accuracy across 10 models and 3 providers (four models hit 100%). 1,300+ LLM evaluations. Zero training.**
1111

1212
Docs: [gcformat.com](https://gcformat.com/) · [Playground](https://gcformat.com/playground.html) · [GCF vs TOON](https://gcformat.com/guide/vs-toon.html)
1313

@@ -179,33 +179,18 @@ Works on dicts, lists, and primitives. Lists of uniform dicts get tabular rows.
179179
| `Session` | Thread-safe tracker for multi-call deduplication |
180180
| `KIND_ABBREV` / `KIND_EXPAND` | Bidirectional kind abbreviation dicts |
181181

182-
## Comprehension Eval
182+
## Benchmarks
183183

184-
Rigorous 3-way benchmark (GCF vs TOON vs JSON) at 500 symbols, 200 edges. 13 structured extraction questions sent to an LLM with zero format instructions:
184+
1,300+ LLM evaluations across 10 models, 3 providers, and 51 independent test runs.
185185

186-
| Format | Accuracy | Tokens | vs JSON |
187-
|--------|----------|--------|---------|
188-
| **GCF** | **100%** (13/13) | **11,090** | **79% fewer** |
189-
| TOON | 92.3% (12/13) | 16,378 | 69% fewer |
190-
| JSON | 76.9% (10/13) | 53,341 | baseline |
186+
| | GCF | TOON | JSON |
187+
|---|---|---|---|
188+
| **Comprehension** (23 runs, 10 models) | **90.5%** | 68.5% | 53.6% |
189+
| **Generation** (28 runs, 9 models) | **5/5** | 1.0/5 | 5.0/5 |
190+
| **Input tokens** (500 symbols) | **11,090** | 16,378 | 53,341 |
191+
| **Output tokens** (100 symbols) | **5,976** | 8,937 | 16,121 |
191192

192-
GCF is the only format with perfect accuracy at scale, at 32% fewer tokens than TOON.
193-
194-
Reproduce: `git clone https://github.com/blackwell-systems/gcf-go && cd gcf-go/eval && GOWORK=off go test -run TestComprehension -v -timeout 0`
195-
196-
## Token Efficiency (TOON's Own Benchmark)
197-
198-
Running [TOON's benchmark harness](https://github.com/blackwell-systems/toon/tree/gcf-comparison) with GCF inserted (their datasets, their tokenizer):
199-
200-
| Track | GCF | TOON | Result |
201-
|-------|-----|------|--------|
202-
| Mixed-structure (nested, semi-uniform) | 170,367 | 227,896 | **GCF 34% smaller** |
203-
| Flat-only (tabular) | 66,029 | 67,837 | **GCF 3% smaller** |
204-
| Semi-uniform event logs | 108,158 | 154,032 | **GCF 42% smaller** |
205-
206-
GCF wins all 6 datasets. On semi-uniform data (the most common real-world pattern), GCF uses 42% fewer tokens than TOON.
207-
208-
Reproduce: `git clone https://github.com/blackwell-systems/toon && cd toon && git checkout gcf-comparison && cd benchmarks && pnpm install && pnpm benchmark:tokens`
193+
GCF wins all 6 datasets on [TOON's own benchmark](https://github.com/blackwell-systems/toon/tree/gcf-comparison). Full results: [gcformat.com/guide/benchmarks](https://gcformat.com/guide/benchmarks.html)
209194

210195
## Links
211196

0 commit comments

Comments
 (0)