Token cost optimization proxy for LLM APIs - intelligent caching, model routing, prompt compression, and usage tracking to reduce costs
-
Updated
Jul 31, 2026 - Python
Token cost optimization proxy for LLM APIs - intelligent caching, model routing, prompt compression, and usage tracking to reduce costs
Token-cheap code-smell CLI for AI agents.
Convert JSON ↔ TOON and generate JSON Schema + compact TOON schema for LLM prompts. Token-efficient, with live token comparison.
MCP-based context engineering patterns (PTC + PTD) achieving 51% token reduction vs ReAct for LLM agent systems
Token-Light, Code-Intensive (TLCI) — A design philosophy for AI agent automation. Use AI only where it's needed. 80-97% cost reduction.
Contextomizer is an ultra-fast, deterministic library for transforming bloated tool outputs, raw APIs, documents, and messy logs into perfectly optimized context for AI Agents 🤖🚀
This monorepo contains the reference implementation and developer tooling for the PowerChain PWRC native-token profile and PowerPay settlement program. It keeps application code, reusable TypeScript packages, Solana programs, canonical IDLs, metadata, environment profiles, validation scripts, and generated release artifacts in explicit boundaries..
Local-first autonomous AI agent with a web GUI and an experimental Token Kernel for bounded context.
Evidence-preserving context compression CLI for agent workflows.
Audit what your output compressor silently deleted: how much of the reduction was real noise, how much was signal. Turns a compressor's 'cuts 90%' claim into two honest numbers.
Route the task before you run it - surface, model and budget in one card. A Claude skill for spending tokens where they earn their keep.
Shared browser interaction schema registry for AI agents. 80-100% token reduction.
Production multi-agent orchestration framework - Claude Code, CLAUDE.md governance, multi-tier model routing, token economics
Public Claude Code plugin marketplace for CULP — reduce Claude Code token usage without reducing output quality.
A lightweight Claude Code hook that intercepts verbose CLI commands and truncates their output before it reaches the model — reducing token usage without sacrificing code quality. Built for Laravel projects but works with any codebase.
Real-time cost tracking and comparison across LLM platforms and models, translating token usage into actual dollars and credits.
Detect, fix, and measure AI coding-agent context-token waste across Claude Code, Codex, Cursor, Gemini & Copilot.
Skill Claude pour decider et executer la conversion JSON <-> TOON (Token-Oriented Object Notation)
Cut LLM costs by 10–30% using better orchestration—not model changes. A visual demo comparing Weft-style pipelines vs map-reduce and full-buffer baselines, showing token usage, cost savings, and % reduction.
Bilingual (中文/English) output-compression skill for coding agents. Benchmarked against caveman and token-diet, and honest about the <1% ceiling.
To associate your repository with the token-optimization topic, visit your repo's landing page and select "manage topics."