Tokenmaxxing

Stop burning tokens on work
CLI tools can do for free

Every file read, search, or edit stuffs full content into context — most of it noise. On a capped plan, that means hitting limits mid-task. The fix isn't a bigger plan; it's running a CLI tool first and handing the LLM only the result.

A single semantic search can consume 24,000+ tokens when an LLM reads the whole file unassisted. The same lookup with the right CLI tool: 188 tokens.

The pattern

Run a deterministic CLI tool first. The tool produces a compact, targeted result — a diff, a filtered file list, an exact passage. The LLM only ever sees that result, not the raw codebase. Token costs drop by 65–99% without changing the answer.

Full file in context
24,437
tokens
→
CLI tool runs
qmd get
deterministic · instant
→
Exact passage
188
tokens

Here are six tools we've benchmarked with real before/after numbers.

Know a tool that should be here? The benchmark is open source — github.com/pmcfadin/agentic-token-bench

Benchmarked tools

Ordered by token reduction, highest first. All runs are deterministic — same input, same output, every time.

Without tool (raw tokens) With tool (reduced tokens)
Y-axis: log scale (token counts vary widely across tools)
Loading benchmark data…

Wire these into your workflow

The benchmark proves the numbers are real. The integration guides show you how to add each tool to Claude Code, Codex, or Gemini CLI in under five minutes — with copy-paste commands and paste-ready CLAUDE.md snippets.

Integration guide → Agent configs →

Know a CLI tool that should be here?

The benchmark is open to any tool that takes file or codebase input and produces a smaller, targeted output. If it's deterministic and saves tokens, it belongs here.

  • Takes file or codebase input
  • Produces smaller, targeted output — a diff, filtered results, or an index
  • Deterministic — same input always yields the same output
Submit a tool →