JevTO
JevTO Helpv0.1.1 · experimental

JevTO / Documentation

JevTO benchmarks: payloads, whole sessions, and limits

A smaller tool response is useful only if the agent can still finish the task. The payload and whole-session measurements answer different questions.

Historical evidence, not a fresh v0.1.1 campaign. The payload record is dated 2026-09-30 and identifies JevTO 0.1.0 and RTK 0.48.0 on Windows. Competitor results apply to that version and setup.

23 generated command scenarios

Each scenario runs a real command in a generated workspace. Tools use their own Claude hook where routed; explicit wrapper cases are labeled in the full report. The dev set was used to tune rules. Holdout and semantic scenarios were written after the initial freeze; the first holdout run exposed bugs that were fixed, so final results are not a pristine unseen evaluation.

Historical totals · estimated tokens = bytes ÷ 4
ArmOutput estimateFacts visible
Native269,73646/46
RTK 0.48.0236,53544/46
JevTO rules16,97945/46
JevTO + Jev13,62846/46

A “fact” is a required string visible in the selected output without recall. This checks selected evidence, not arbitrary semantic correctness. The report also charges the full native payload when a fact is missed; that is a benchmark accounting rule, not observed agent behavior or a provider bill.

RTK is tighter on some Rust/Go failure outputs and tiny passing runs. Rules miss the vocabulary-mismatch search answer. The Jev arm recovers that answer in this small set. The source includes negative results and raw views.

Whole-agent pilots

Codex CLI 0.158.0-alpha.2.1, gpt-6-sol: one parser task, four arms, three runs each. All 12 sessions passed the hidden holdout. Median input tokens were 198,520 native and 180,533 with JevTO (about 9% lower). This small result does not establish broad savings.

Claude Code 2.1.283, Haiku 4.5: all 12 sessions passed visible tests and failed the same hidden holdout. These are not verified successful task completions. Reported list-price estimates are not bills.

A separate Cursor pair had unequal holdout outcomes and cannot support a successful-task savings comparison. See the public README's session evidence and raw records.

Reproduce or challenge the results

From a checkout of the public repository with the relevant scenario toolchains installed:

Terminal
cargo build --release --locked -p jevto
python benchmarks/token_bench.py
python benchmarks/render_charts.py

For the generated holdout suite, use python benchmarks/token_bench.py --suite holdout. Review the harness and its prerequisites before running it. Live adaptive requests are a separate opt-in using --live-jev and your runtime OpenRouter key; those requests can incur provider charges. No paid campaign runs merely from browsing this site.

Cached Jev decisions are committed for inspection and replay. Historical cost fields predate the newer request-accounting corrections; they are not complete bills. The fresh-task note distinguishes offline evaluator controls from real coding-agent results.

What you can conclude

JevTO can substantially reduce selected command outputs and preserve exact local recall on tested captures. The current evidence does not prove universal fact preservation, lower bills across projects, or unchanged task quality. Try it on your own verifier and report both successful and failed runs.