通过将文本上下文渲染为图像来减少 Claude 代码令牌的使用
Cut Claude Code's input tokens by rendering bulky context as images — the same system prompt, tool docs, and history, in a fraction of the tokens.
An image's token cost is fixed by its pixel dimensions, not by how much text
is inside it. Dense content (code, JSON, tool output) packs ~3.1 chars per
image-token vs ~1 char per text-token on real Claude Code traffic. The
reader is the same vision channel that Anthropic's computer use already
relies on for screenshots. pxpipe is a local proxy that uses that channel
for context: it rewrites the bulky parts of each request into compact PNGs
before it leaves your machine. At current Fable
list prices that lands as a ~59–70% lower end-to-end bill — but prices
move and workloads differ, so the durable number is the token cut itself,
measured per-request against a free count_tokens counterfactual in
~/.pxpipe/events.jsonl.
This is what the model sees instead of text:
~48k chars of system prompt + tool docs: ≈25k tokens as text, ≈2.7k image tokens as this page. Real pipeline output; the model reads renders like this at 100/100 (see benchmarks).
*Eight years of context growth, in characters. Every text line tops out near
~4M chars (a 1M-token window at ~4 chars/token); Grok 4.5 is shown as a
text-window point only (500K). The orange overlays are the same 1M
windows read through pxpipe images — ~19.0M chars for Fable 5 (4.8×) and ~21.3M chars for Gemini 3.6 Flash (5.3× text capacity). Density is measured from a live render at
generation time, not hand-typed: regenerate with
npx tsx scripts/gen-context-chart.ts
(source).*
Fable 5 (the default, 100/100 reader) — plain left, pxpipe right:
https://github.com/user-attachments/assets/1c8ee63a-fcd7-4958-917b-da788d718349
pxpipe counts an exact token 10/10 across 39 imaged filler files
(matches grep line-for-line), gets the multi-step ledger arithmetic right,
and ends the session at $6.06 with context to spare (73.5k/1M) vs
$42.21 at 96% full. One caveat visible in the clip: the pxpipe arm
needed a nudge to match the requested one-line output format.
npx pxpipe-proxy # proxy on 127.0.0.1:47821
ANTHROPIC_BASE_URL=http://127.0.0.1:47821 claude # point Claude Code at it
Dashboard at http://127.0.0.1:47821/: tokens saved, every text→image conversion side by side, kill switch, live model chips. Responses stream normally — pxpipe compresses the request only, never the model's output. Recent turns stay text; the system prompt, tool docs, and older bulk history are imaged.
pxpipe warppxpipe warp -- claude # also: cursor-agent, codex, or a shell alias
Same thing without ANTHROPIC_BASE_URL, so /remote-control, claude.ai
connectors, and first-party gates keep working. Full instructions in the
dashboard.
api.anthropic.com/v1/messages is routed by default. Agents that reach their
provider over some other base URL need a rule for it, and a rule that names a
port matches only that port:
pxpipe warp --route '127.0.0.1:9090/v1/*=http://127.0.0.1:47821' -- codex
You can render text, files, or diffs to PNG pages without running the proxy or connecting Claude Code:
npx pxpipe-proxy export src/
cat prompt.txt | npx pxpipe-proxy export --stdin
npx pxpipe-proxy export --git
If the package is installed, use pxpipe export instead of
npx pxpipe-proxy export.
Each run writes a fresh pxpipe-export-XXXXXX/ output folder (the exact path
is printed when the command finishes) containing page-*.png, factsheet.txt,
manifest.json, and prompt.txt. Upload the PNG pages and paste the prompt
into image-upload clients such as Cursor when you want dense visual context
without running the proxy.
CLAUDE_CODE_SUBAGENT_MODEL=claude-sonnet-4-6, or model: sonnet in
agent frontmatter).eval/./anthropic/messages and typically lands ~60–70%. Details and measured
splits: docs/CACHING_AND_SAVINGS.md.claude-opus-5: weaker recall than Fable 5 (verbatim 2/15 vs 13/15), good
enough otherwise (100/100 arithmetic, 0/16 never-stated), ~4.7× context before
/compact. Suggested effort: medium. Details: FINDINGS.md.PXPIPE_MODELS=claude-fable-5,gemini. The gemini
base covers every Gemini id (3.6/3.7/3.8 Flash, Pro, 4, 5, and future
versions); to opt Gemini out, drop gemini from PXPIPE_MODELS or click the
chip off. Opus 5, Sol, GPT 5.5, and Grok are opt-in only (dashboard chips or
PXPIPE_MODELS). The exact Sol id still matters. Sibling variants such as
gpt-5.6-terra do not
inherit Sol's allowlist or render profile. PXPIPE_MODELS=off disables
imaging. Everything else passes through byte-identical. On the GPT path,
tool definitions stay native JSON and no Anthropic cache_control
markers are used. Responses history compression recognizes completed
function_call/function_call_output pairs, including OpenCode's parallel
calls-then-outputs rounds: only old closed rounds are imaged atomically;
every open call and malformed/orphan state remains native. The base profile
keeps the newest six completed pairs and allows 32 images; Sol keeps one pair
and allows 64 images, while Grok allows 24 images. Opt-in long-session
coverage can be changed (defensive cap 100) with
PXPIPE_GPT_HISTORY_MAX_IMAGES=48 after validating the provider's request cap.gpt-5.6-sol and Grok use native 14px
JetBrains Mono glyphs in a 9×16 cell, 84 columns, and a 764px full-width
strip; Claude keeps its 312-column, 1568×728 5×8 Spleen profile. These
are selected by exact model id, including history pages and profitability
math. Recognized IDs can ride in the bounded factsheet, and
recent/open tool state stays native.
Sol receipts and
profile evidence.PXPIPE_MODELS=claude-fable-5,grok-4.6 or the dashboard chip.
eval/grok-density/QUALITY_RESULTS.md.This matrix shows coverage as well as scores. — means the model was not run
on that test; it does not mean zero. Arithmetic uses novel random-number
problems. Gist, state, and never-stated probes share one corpus. Never-stated
is confabulations, so lower is better. The numbers at column is the render
geometry the row's scores were measured at; a model's shipped profile can
differ (Sol and Qwen ship the measured 14px/84 geometry, but their broad-suite
numbers predate it).
| model | numbers at | arithmetic (N=100) | gist (N=98) | state (N=18) | never-stated (N=16) | dense hex (N=15) | profile provenance and receipts |
|---|---|---|---|---|---|---|---|
claude-fable-5 |
Spleen 5×8, 312 cols (shipped) | 100/100 | 98/98 | 18/18 | 0/16 | 13/15 | June 2026 production profiles: arithmetic + hex, gist/state/guards |
claude-fable-5-1 |
Spleen 5×8, 312 cols (Fable 5 profile) | 100/100 | 95/98 | 18/18 | 0/16 | 6/15 | Fable 5 profile, no geometry of its own; 3 gist misses are image-arm negation flags answered UNKNOWN (0 confabs). Same-day Fable 5 control on the identical harness/PNGs reproduced 100/100 arithmetic and 30/30 tier-2 gist, so the gist/hex gap is the model, not the harness (hex control not rerun): arithmetic, dense hex, gist/state/guards |
google/gemini-3.6-flash, 3.7-flash |
Spleen 5×8, 312 cols (shipped) | 100/100 | 98/98 | 18/18 | 0/16 | 14/15 | current shipped profile: quality results |
claude-opus-5 |
Spleen 5×8, 312 cols (shipped) | 100/100 | 94/98 | 17/18 | 0/16 | 2/15 | current profile: arithmetic, gist/state/guards, dense hex |
gpt-5.6-sol |
Spleen 5×8, 152 cols; ships 14px/84 | 98/100 | 83/98 | 17/18 | 4/16 | 0/15 | broad suite predates the shipped 14px profile; 14px pilot: 7/8 exact, 0 inventions, gist/guard pass: pilot |
claude-opus-4-8 |
Spleen 5×8, 312 cols (historical) | 93/100 | 77/98 | 18/18 | 0/16 | 0/15 | historical profile: arithmetic, gist/state/guards, dense hex |
grok-4.5 |
JetBrains Mono 14px, 84 cols (shipped) | 100/100 | 97/98 | 17/18 | 0/16 | 0/15 | native 14px/84 quality suite (live profile); quality, native-sweep |
grok-4.6 high |
JetBrains Mono 14px, 84 cols (shipped) | 100/100 | 97/98 | 17/18 | 0/16 | 0/15 | native 14px/84, reasoning high; quality |
moonshotai/kimi-k3 |
Spleen 5×8, 152 cols (generic default) | 79/100 | 84/98 | 15/18 | 1/16 | 0/15 | generic GPT profile, no measured geometry of its own: quality results |
qwen-3.8 (@cf/qwen/qwen3.8-27b) |
Spleen 5×8, 152 cols; ships 14px/84 | 98/100 | 72/98 | 11/18 | 0/16 | 0/15 | broad suite predates the shipped 14px profile; 14px pilot: 8/8 exact, 0 inventions, 11/15 hex: pilot & quality |
glm-5.3-flash (@cf/zai-org/glm-5.3-flash) |
Spleen 5×8, 152 cols (default fallback, nothing shipped) | 36/100 | 57/98 | 6/18 | 0/16 | 0/15 | 5×8 is illegible to GLM (0/15 hex); 14px pilot: 10/15 hex, all misses single-glyph confabs, guards 0/16: pilot & quality |
Offline export of the same deterministic 454,045-character dense record corpus through each complete profile produced:
| model profile | pages | text estimate | image tokens | savings |
|---|---|---|---|---|
| Claude, Spleen 5×8 | 17 | 122,715 | 23,856 | 80.6% |
| Sol, JetBrains Mono 14px | 45 | 122 |
暂无开放 Issues,或尚未同步最近议题。