Measured 2026-08-17
Fewer tokens. Same answers.
This page is the results. Not a product tour. We ran the local replay harness on synthetic tool output and published every number it produced. Mixed real sessions save less. Source files are left alone on purpose.
Key results
Headline figures from the 2026-08-17 harness run, default settings. The 90% average is the noisy fixtures only. The table below also includes the source-file row that saved nothing.
-
90%
avg. saved on noisy fixtures
-
48,453
tokens cut from 51,970
-
4/4
JSON error facts kept
-
0%
source files rewritten
Token savings by scenario
Each row is one request body the harness compressed. Original and after are estimated input tokens of the whole request, the same estimator the app uses. Percents are rounded.
| Scenario | Original | After | Saved | Ratio |
|---|---|---|---|---|
| Code search 100 JSON results | 5,524 | 574 | 4,950 | 90% |
| CI log cargo test spam | 12,714 | 160 | 12,554 | 99% |
| GitHub issue 80 comments | 5,913 | 707 | 5,206 | 88% |
| PowerShell listing Get-ChildItem recurse | 3,452 | 586 | 2,866 | 83% |
| Multi-tool agent JSON + log + HTML in one turn | 21,222 | 1,032 | 20,190 | 95% |
| HTML docs fetch page with scripts and nav | 3,145 | 458 | 2,687 | 85% |
| Source file 220-line module read | 3,311 | 3,311 | 0 | 0% |
| Total | 55,281 | 6,828 | 48,453 | 88% |
How to read the table
The high ratios are the bulky tool payloads: search JSON, test logs, issue threads, directory listings, fetched docs. The 0% row is not a miss. Source files already have to stay intact, so the engine leaves them alone.
- Code search (90%). 100 JSON search hits. Rank, URL, title, and snippet repeated a hundred times. The distinctive middle result still resolves after compression.
- CI log (99%). Hundreds of PASS lines plus one ERROR. The failure line is what you need; the rest is spam that would ride along on every later turn.
- GitHub issue (88%). 80 comment objects in one JSON array. Thread padding compresses; the shape of the thread remains.
- PowerShell listing (83%). A recursive Get-ChildItem dump with progress noise. The access-denied line stays readable.
- Multi-tool agent (95%). Search JSON, a test log, and an HTML page in a single turn. This is the closest fixture to a busy agent loop.
- HTML docs fetch (85%). A page with scripts, styles, and nav around a short article. Chrome goes; article words stay.
- Source file (0%). A 220-line module read. 3,311 tokens in, 3,311 tokens out. Code is not a compression target.
Quality checks
Cutting tokens is easy if you are willing to lose the answer. These checks are the ones we can run without calling a model: facts that must survive, pages that must still be readable, and blocks that must not change once they have been sent.
-
4/4
JSON retrieval
Error, code, resolution, and affected count at log index 67 kept inline.
-
100%
HTML article recall
Every article word survived extraction. Scripts and styles were dropped.
-
3/3
Needles recovered
Facts buried in compacted JSON, logs, and HTML either stay inline or resolve from the local originals store. The model can pull any stub back.
-
Stable
Prompt-cache prefix
Already-sent compressed blocks replay the same bytes on the next turn.
-
Untouched
Source files
Code passes through. We compact tool noise, not your repo.
JSON retrieval
100 production-style log objects. Index 67 is an error. Four fields must remain readable in the compressed request, without a round-trip to disk: the error string, the code, the resolution, and the affected count. The harness kept 4/4.
Compact JSON is only useful if the unusual row is still there. Repeating heartbeats can go. The one error cannot.
HTML article recall
Word overlap of the article body after scripts and styles are dropped. 100% recall means every article word survived. Nav, analytics, and CSS were never the answer.
This is not a third-party extraction F1. We did not run Scrapinghub, trafilatura, or any published HTML dataset. Recall against our own article fixture is the claim. Precision is lower on purpose: leftover chrome that is not in the article-only ground truth still sits in the extracted text, and we do not dress that up as a leaderboard score.
Needles recovered
A distinctive string is buried in compacted JSON, in a noisy log, and in the HTML article. After compression it must still be present in the request, or retrievable from the original that stayed on disk. The harness recovered 3/3.
That is the reversible contract: trim the noise in what the model is billed for, keep the full payload on the machine so a buried fact can come back if the model asks.
Prompt-cache prefix
Providers discount tokens that match a previous prefix. If a compressed block that already went out is rewritten on the next turn, that discount is gone and you pay again for history you already sent.
The harness compresses a tool result, then sends a follow-up turn that still contains that result. The already-sent block must replay as the same bytes. It did. We do not publish a dollar figure for that discount. We publish that we did not bust it.
Limitations
An honest results page says where the engine does nothing, and where the fixture average will not match a real week.
What we do not compress
- Source files. Reads of your code pass through. The 220-line module in the table is unchanged, 3,311 tokens in and out.
- Short turns. Tiny messages are not worth touching. The savings live in bulky tool output, not in "ok, try that."
- Already-compact output. Tight structured dumps (a short grep, a small schema) often have nowhere to go.
- Your prompts and the system prompt. We trim tool noise in the history, not the instructions you typed.
- Images. Picture payloads are not treated as text to squash. They are left as the client sent them.
When it helps
The fixtures that moved are the ones that look like a long agent afternoon: tool output that is large, repetitive, and still sitting in history on turn 20.
- JSON-heavy turns (search hits, API lists, log arrays). The 100-item search fixture saved 90%.
- Build and test output. The cargo-test spam fixture saved 99%. One ERROR line is the payload; the PASS flood is not.
- Fetched HTML with scripts and nav around a short answer. The docs fixture saved 85% and kept every article word.
- Multi-tool loops that pile several of those into one turn. The combined fixture saved 95%.
Those tokens are not billed once. They ride along until the session ends. That is why a noisy hour hurts the weekly cap, not just the five-hour window. More on that at How it works.
When it adds little
- Code-only sessions. If the agent is mostly reading and writing files, the table's 0% row is your week.
- Short conversational exchanges. A few sentences of back-and-forth have almost nothing to trim.
- Already-small tool results. A tight grep or a tiny schema is already compact.
- A mixed real week. Production sessions blend all of the above. They will not look like the 90% fixture average. The free audit on your own transcripts is the number that matters for you.
What we do not claim
Other pages in this category publish fleet telemetry, proxy latency percentiles, and LLM accuracy evals. We have those kinds of numbers only if we measured them. We did not, so they are not here.
- Not production telemetry. These rows are synthetic fixtures, not a fleet median across live installs. We do not invent a "tokens saved" total for everyone.
- Not an LLM accuracy eval. We have not run SQuAD, HotpotQA, or live agent grading. If a vendor quotes those, they called a model. We are publishing what the engine itself can prove.
- Not a latency table. We did not measure round-trip overhead for this page. No millisecond column, no P50/P99.
- Not a promise about your bill. The 90% average is the noisy fixtures. Your mix is the audit.
How we measure
The harness builds realistic request bodies (search JSON, test logs, issue comments, PowerShell listings, fetched HTML, a source file) and runs them through the same compression the app uses, on default settings.
Tokens and ratio
Tokens are an input-token estimate of the whole request, before and after. The estimator is the one the app uses, so the table matches what you would see in product.
ratio = 1 - (tokens_after / tokens_before) A 90% ratio means the after size is 10% of the original. Percents on this page are rounded; the token counts are exact from the 2026-08-17 run.
The noisy average (90%) is the mean of the six rows that saved anything. The table total (88%) includes the source-file passthrough, which is the fairer "all fixtures" figure.
Quality checks
- JSON facts. One error object at index 67 in a 100-entry log. Four fields must remain readable without fetching the original.
- HTML recall. recall = |predicted ∩ article| / |article|. 100% means every article word is still in the extracted text. We do not headline F1.
- Needles. A distinctive string buried in JSON, logs, and HTML must remain in the request or resolve from the local original.
- Prefix stability. A compressed block that already went out must replay as the same bytes on the next turn.
- Code passthrough. A 220-line module must not shrink. The published row is 0%.
The published floors are locked in CI against this harness. If the engine regresses, the site numbers fail the build. We do not ship a public eval suite to clone and re-run; the audit on your own transcripts is the reproduction that matters.
See it on your own transcripts
The fixtures above are noisy on purpose. The free audit reads local Claude Code and Codex logs and shows where your week actually went. Nothing uploads.