doc.html

~ the evidence — what it does, where it doesn't fit ~ every figure read out of the public record

what it does, where it doesn't fit

The claims this site makes, with the numbers behind them and the bounds that scope them. A bound travels inside its claim, the way ± travels with a measurement — a saving that inverted on another harness says so in the same breath, a recall that cost extra tokens carries the cost. Every figure on this page is read out of the public record, a witnessed doc.html these paragraphs link into, and the losses are in it at full weight: the complete ledger — every verdict, including the runs that went against us. Nothing on this page was measured here; the record governs.

at size — the document outruns the reader, and stays readable

The largest sealed artifact in the record is one 72.5 MB doc.html of 17,631 sections, far past any context window. In a sealed live run across Claude, GPT and Kimi routes, models read their way through it and cited their way back: all 480 of 480 navigation turns completed (mean wall clock ran roughly 8 to 34 seconds by vendor and mode), and every one of 240 self-citations checked out against the bytes it named. [ apparatus memori · PR #48 ]

The reading was reading, not recall: a canary token and a fabricated gloss existed only inside the emitted artifact, and the models touched them in 120 of 120 contacts — 116 of 120 under the stricter byte-for-byte discriminator, the four misses (two case-changed thetas, two double-escaped decimal entities) dissected to the codepoint in the record rather than rounded away. The flat orientation layer kept a real tail at that size: the worst single lookup, a canary on the Kimi route, took 209.3 seconds. [ apparatus memori · PR #48 ]

what it costs — where selective reading pays, and where it doesn't

On a document big enough to matter, reading the map beats handing over the pile: a 759 KB corpus over 63 graded cells cut median effective tokens by 52.9% — bytes loaded fell 65.5% and mean answer quality rose 1.19 in the same run — and the bound rides with the number: the saving did not survive a change of harness. Same model, same corpus, same questions on another harness, and the preregistered stability condition failed, the figure inverting to +135.7% from −16.9%. The run identity governs, not the format. [ baseline cost · PR #1 · cross-harness · PR #8 ]

Against vector RAG it is not the cheap option: on a 7.80 MB, 619-section document against vanilla Chroma with Titan Text Embeddings v2, walking the document spent 20–50× more effective tokens (roughly 677K against 20K). What the spend bought was evidence — median evidence recall of 0.875 on the overflow questions against RAG's 0.25 at top-k=8 — with hand-authored summaries on 32 likely cited sections part of the condition, and no claim made against stronger RAG systems. [ rag comparison · PR #13 ]

On a small file there is nothing to win: a roughly 3 KB document ran +60.3% effective input tokens with answer quality flat, and a later corrected sweep scored every arm alike — a final null. Selective reading is for documents big enough to outgrow a context window; under a few KB, hand the model the file. [ tracer #11 · PR #34 ]

what a machine can check

The file does not check itself, and no model checked it unprompted: 0 of 27 spontaneous hash verifications across Claude, GPT and Kimi. Prompted to hash by hand, aggregate accuracy was 10/17, about 59%, with one shared failure — the wrong canonical byte slice. Put a deterministic reader in the loop and the check lands: 27 of 27, the models invoking the helper and relaying its verdict correctly in every cell. The integrity property is operational with a verifier, not by disposition. [ drift defense · PR #11 ]

A mechanical judge holds the writers to the bytes: across 360 live cells it found eight real failures in 206 — six non-contained quotations, two witness mismatches — and caught all eight. Writer survivability came to 358/360 (0.9944), and the residue leg's 0/40 carries a Wilson 95% upper bound of 8.8% rather than a true zero. [ apparatus criticus · PR #44 ]

a format, or one program

A stranger can build a reader from the written spec and agree with the shipped ones: a blind C#/.NET subject matched the expected verdict on every case of a twelve-case battery across four readers. One raw-text cell split the four and is carried as named implementation debt; blindness was instruction-isolated and self-attested, the sample small, the run read-side only. [ reader convergence · PR #72 ]

Two readers ship because one was not enough: a forged witness planted in a decoy attribute was accepted by an attribute-reading path in the Python reader and refused by the JavaScript one. The forgiving tokenizer was deleted rather than patched, and the pair re-sealed as a unit with both readers' shipped bytes pinned — two readers, agreeing by refusal. [ reader repair · PR #134 ]

what travels in the file

Memory travels: eight invented project facts, hidden from the manifest summaries, were recovered 24 of 24 by a model given the document, while the no-document control recovered 0 of 24 — at a cost that rides with the claim: on that 118 KB corpus, selective recall spent +38.7% effective tokens over a capable whole-document reader, which also reached 24/24. The recall held; the economy did not. [ memory + steering · PR #85 ]

Steering travels the same way, and placement decides whether it governs: the same non-default authoring rule complied 3 of 3 when always loaded, 0 of 3 when hidden under a broad body heading, and returned to 3 of 3 under a task-obvious heading. [ memory + steering · PR #85 ]

The thread survives a change of hands: a second harness given no document path, no memory path and no hidden state followed the links in the grown chat artifact and made all five required identifications — single-sample, builder-evaluated, nonblind; an existence proof, not a success rate. A 28-turn live conversation kept as witnessed turns passed the same kind of proof, and the shape it produced is now specified; long-chat token savings are not yet measured, and the document does not perform the model call. [ trinity handoff · PR #39 · chat v3 · PR #91 ]

A shelf of linked pages answered the ordinary-retrieval questions as well as the markdown wiki it was ported from — 12/12 on both sides — and localized knowledge-layer corruption the comparison wiki's own health tooling missed. The preregistered reader leg is inconclusive (its vacuity guard fired), and the cross-file verifier is two-level, not recursive. [ multi-file wiki · PR #87 ]

how the format got its shape

Some of the format's furniture exists because an early hypothesis broke, and the break is in the record.

Nested indexes exist because the flat map broke first: a thin flat mirror of a 619-section corpus came to an estimated 74,446 tokens against a 12,000-token reading budget — 6.2× over — failing in staging before an agent was ever invoked. That is why large documents fold their maps. [ flat mirror · PR #29 ]

The deterministic readers exist because disposition failed: the 0-of-27 above is why verification is a command that ships with the format instead of a hope about model behavior. [ drift defense · PR #11 ]

The pitch at documents big enough to matter exists because the small-file run went null. The format starts paying where the context window stops. [ tracer #11 · PR #34 ]

And one edge is simply open: deep-hierarchy selection with a live selector has not met its preregistered thresholds — completion 227/264 (85.98%) against a 90% rule, grade 0.50 against 0.80, with zero answered-but-wrong across 1,344 live cells. The capacity walls held; the thresholds did not. The site claims nothing here. [ bounded return run-02 · PR #114 ]

[ ↑ top ]

the record this is read from

The public evidence document is the source for every figure on this page — a doc.html itself, manifest at the top, each section pinned to its own bytes. It carries what a summary has to compress: which harness, which model family, how the grading was done, what the seal covered, and which later run replaced an earlier headline. A figure here is a pointer to the section that carries the detail, and that section governs.

the laboratory paths, pull-request numbers and seal tags printed in the receipts column are provenance identifiers from the upstream repository. they are not public links, and the raw populations are not shipped.

git clone https://github.com/NDOTO-G/doc.html
cd doc.html
node tools/verify.mjs documents/evidence.doc.html
python tools/verify.py documents/evidence.doc.html

a PASS means the manifest and the body agree on every witnessed byte span. it does not rerun the live-model studies.

read it whole, in the format it argues for:

[ ↑ top ]