JevTO Benchmarks v0.1.1 · EXPERIMENTAL
JevTO: output selection for coding agentsExperimental · v0.1.1Experimental
JevTO logo

Less tool output.
Exact recall.

Give Claude Code and other coding agents a smaller view of command output. JevTO selects real lines, preserves exit codes, and stores the original bytes locally for recall. Start with rules only; no API key needed.

Command output
−94%
rules-only estimate, 23 scenarios
Required facts visible
45/46
without a recall
Session input, Codex
−9%
median, n = 3, equal outcomes

Free · Apache-2.0 · experimental v0.1.1. Payload reduction is not measured bill savings. Read the evidence.

Output inspector

Same command. Three views.

REAL BENCHMARK CAPTURES
exit 1

Agent's goal: Fix the failing parser test

Loading capture…
Exact text each tool's Claude Code hook delivered in the benchmark run. Temp paths shortened to <tmp>.
Deterministic mode: no network, no accountRecall returns historical bytes; it never reruns a command
How JevTO selects outputMethod

Keep a useful view. Recall the rest.

Passing-test inventories, progress redraws, and regenerated lockfiles can fill an agent's context. JevTO selects output line by line and records why it defers a section. Selection can miss useful evidence, so every view includes a way to recover the captured bytes.

  1. 1

    Capture exactly

    The command runs with your permissions. Stdout, stderr, and the exit code are stored separately and hashed in a local store.

  2. 2

    Protect recognized failures

    Recognized failures, assertions, panics, warnings, and exit status are protected by local rules. Jev ranking cannot override those protections.

  3. 3

    Fold the ceremony

    Recognizers for Rust, Go, Python, Node, Jest/Vitest, and TAP test inventories, Cargo progress, lockfile and minified diffs, and long logs replace noise with a count and a gap marker.

  4. 4

    Recall on demand

    Each gap names a section. jevto recall ID --section … returns those bytes exactly. If the view would not be smaller, the original is delivered unchanged.

What the command printed
test_accepts_equals ... ok
test_accepts_tabs ... ok
… 218 more ok lines …
test_rejects_leading_space ... FAIL
AssertionError: '7' is not None
Ran 221 tests in 0.021s
FAILED (failures=1)
What the agent reads
jevto exit=1 omitted=225
... 220 passing tests hidden [stderr-0-13420]
test_rejects_leading_space ... FAIL
AssertionError: '7' is not None
Ran 221 tests in 0.014s
FAILED (failures=1)
recall: jevto recall d3768e28
3,572 → 200 tokens on this run. The gap marker keeps the count, so the agent still knows 220 tests passed.
Optional: Jev, the decision layer.

Jev can rank search results, diffs, and long-output sections by meaning. The default auto mode enables this when an OpenRouter key, a goal, and eligible output are present. Bounded snippets or line-shape digests and the goal may be sent to OpenRouter (≤ 64 KiB, 8 s timeout, secret filtering). Use JEVTO_MODE=rules for no network calls. The historical live payload run kept 46/46 benchmark facts visible; rules kept 45/46. Understand what leaves your machine.

Capture → Protect → Fold → RecallEvery decision is written to a local receipt
BenchmarksReproducible

Fewer tokens only count if the answer survives.

23 generated scenarios cover tuned dev cases, a holdout, and vocabulary-mismatch cases. Each tool uses its own hook where routed; explicit-wrapper cases are labeled in the raw report. Tokens are bytes ÷ 4 for one command's output. A fact is a required string visible without recall, such as the failing test or relevant commit. The final holdout results include fixes made after its first run.

Historical rules-only payload totals by suite. Native: 269,736 estimated tokens, 46/46 facts. RTK 0.48.0: 236,535, 44/46. JevTO rules: 16,979, 45/46.
Per-scenario dot plot on a log scale. RTK is smallest on Rust and Go failures and the tiny passing run, but does not route Python, Node, or log output, and misses the needed fact on git log and the 14-file diff. JevTO misses the fact on the vocabulary-mismatch search question.
Log scale. Hollow red-marked dots lost a required fact. RTK's hook doesn't route Python, Node, or plain log dumps, so those rows sit at native.

Historical results: 2026-09-30 on Windows, JevTO 0.1.0 and RTK 0.48.0. These are not a new v0.1.1 benchmark or current-competitor claims. Method, limitations, and reproduction · Raw views and JSON.

honest-reading.txtLimits

What the numbers do and don't say.

Where it wins

  • Test runners RTK doesn't route (Python, Node) and long logs: −69% to −98% with every fact kept, on scenarios written after the rules froze too.
  • Lockfile-heavy diffs: 11,484 → 120 tokens; a 14-file mechanical refactor shows the edit once and the real fix in full.
  • Keeps facts that truncation drops: the git log commit and the multi-file fix line.
  • The small Codex pilot used fewer median input tokens with 12/12 hidden-holdout passes across its four arms.

Where it doesn't

  • RTK is tighter on Rust and Go failures and tiny passing runs.
  • Rules missed one vocabulary-mismatch search answer. The historical live Jev arm recovered it, but that does not establish universal answer preservation.
  • The Claude pilot failed every hidden holdout. Smaller payloads and list-price estimates do not establish lower bills or successful tasks.
  • n = 3 on one small task per agent. A Cursor pair had unequal outcomes. Indicative, not significant.

The benchmark scenarios were written by the JevTO author. They're generated from scratch by a committed script so anyone can rerun them, or add scenarios that JevTO handles badly.

Supported agentsProperties

The route is explicit.

Every installer previews first, applies only with --apply, backs up what it changes, and removes only its own entry. A failed hook leaves the command native. The Claude hook never sets a permission decision.

HostRouteStatus
Claude Codeinit-claude-auto: rewrites simple test, build, lint, search, and git diff/log/show commands; records the prompt as the session goalLive sessions
Codexinit-codex-pre-hook: PowerShell commands on the same allowlist (cargo, rg -n -H, python -m unittest, …), Windows onlyLive sessions
CursorProject rule for one exact verifier; opt-in exact-command MCP runnerNarrow
Any MCP hostRead-only jevto_recall, jevto_review, jevto_statusImplemented
Any shelljevto run -- PROGRAM ARGS…Implemented

Setup, tested versions, and removal instructions. jevto doctor --json reports coverage and mode without making a network call.

Local storagePrivacy

Local capture. Optional remote ranking.

  • Deterministic mode makes no network requests and sends no telemetry.
  • Captures expire after 24 hours; the store is capped at 1 GiB.
  • Command arguments and environment variables aren't stored.
  • In default auto mode, an OpenRouter key enables eligible remote ranking from the current workspace. Set JEVTO_MODE=rules to stay local.

The store isn't encrypted at rest, and captured output can contain secrets. Read storage, network, and removal details.

Command PromptRust stable

One binary. Start with one command.

Download v0.1.1 for Windows x86_64, Linux x86_64, or macOS Intel/Apple silicon. Every archive has a checksum. No JevTO account is needed.

C:\src\jevto>
cargo install --git https://github.com/mattbaconz/jevto --tag v0.1.1 --locked jevto
jevto doctor

# Claude Code: preview, then apply
jevto init-claude-auto --workspace .
jevto init-claude-auto --workspace . --apply

# Any agent or shell
jevto run --mode rules -- cargo test --workspace
jevto recall last
jevto gain