Optimize token usage for Claude API calls
-
Updated
Aug 17, 2026 - JavaScript
Optimize token usage for Claude API calls
Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.
Measure token savings per AI coding agent, optimize context, and share a live local knowledge graph across 16 CLI clients.
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser agent-to-agent handoffs (−53%, lossless in benchmarks) with no loss of quality. Works with Claude Code, Cursor, GitHub Copilot, Codex, Gemini CLI, Windsurf, Cline & Kiro.
Claude Code plugin that tracks token usage, identifies wasted context, and saves 30-50% on API costs. Heatmaps, ROI reports, budget alerts, efficiency scores, git-aware suggestions — all local, zero config.
The social network for AI agents. Paste one prompt and your agent writes your profile, posts what you ship and finds people who build what you build — you approve. Plus E2E-encrypted agent rooms, a walkable code town, a 3D particle wallpaper, and token filters for Claude Code, Codex & Cursor. MIT CLI + SDK.
EGC gives every AI coding agent the same brain. Shared memory, skills, and live context across Cursor, Claude Code, Copilot, Aider, and 20+ AI coding tools with zero configuration. Every tab, terminal, and AI stays automatically synchronized. One brain. Everywhere.
Measured, monotonic, lossless input savings for Codex.
Local Claude Code and Codex plugin that bounds oversized and repeated tool output, preserving complete redacted artifacts. Runs locally, makes no LLM calls.
Claude Code skills for developers who code like cats — never more effort than the problem requires.
Quieter sessions for Claude Code: less narration, shorter tool output and concise answers focused on the result.
Eco mode for Claude Code. /eco: -31% to -73% output tokens with critical findings intact; /eco-max: up to -75% with lowered effort. Measured hardest on Claude Fable 5 (fable5), deep-studied on Sonnet 5, works on Opus 4.8 too. We publish our negative results. 82 raw benchmark runs.
Restore prior Claude Code AND Codex sessions with zero LLM calls. Cross-tool session handoff, cache-expiry prevention, real cost dashboard. One plugin, both hosts.
Chrome Extension that lets you continue any AI conversation anywhere—without losing context. No more copy-paste—move full context across AI tools.
Verdict-first output for AI coding agents. Tiny prompt + installer for Claude Code, Codex, Gemini, Cursor, opencode, and 30+ agents.
Deprecated — Catalyst local runtime, kept for history. Use coalesce-labs/catalyst-dev-skills and coalesce-labs/catalyst-cloud-skills.
Layered token-optimization pipeline for DeepSeek Harness: output ladder, MCP lazy loading,compaction driver, cache-hit reporting. Built on real DSH plugin APIs; ~40-60% input saved in long sessions.
Local-first context compression for AI coding tools. One binary saves 85-93% of redundant tokens across every LLM call.
AI-powered MCP proxy for Playwright and Figma. Optimize Claude Code and AI agents with session recovery, context reduction and intelligent tool routing.
Free desktop app: run Claude Code with multiple accounts and auto-switch when one hits its 5-hour limit. 24/7 auto-loop, parallel chats, live cooldown timers. No API key needed.
To associate your repository with the token-optimization topic, visit your repo's landing page and select "manage topics."