You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Rolling context compression for Claude Code: a zero-dependency proxy that summarizes old messages and keeps recent turns verbatim — beat the context-window limit and cut token & cache costs. Turn-aware keep policy, concurrency-hardened.
Unified compression pipeline for LLM inputs: trim irrelevant rows and fields, re-encode the rest in the cheapest lossless format (TOON / JSON / CSV, measured on your tokenizer), and account for every token saved with an audit trail of what was removed. 95% fewer tokens on realistic payloads, 100% needle recall.
A research-oriented framework for adaptive context engineering in RAG systems, exploring structured representation, compression, and progressive context expansion through controlled experiments.RAG
Up to 99% token savings on AI agent context. Free. Provider-agnostic (Claude, Cursor, Codex, Cline, any MCP agent). 91.62% over 626.8M auditable tokens. Try: nuxs.ai/playground
Local-first AI computer control plane: governed model routing, durable memory, context compression, bounded agents, MCP, and receipt-backed execution on Windows and network AI boxes.
TokenPack packs long documents, codebases, PDFs, and folders into compact, evidence-dense LLM context using local embeddings, evidence scoring, and budget-aware selection.
LLM token compression via images: a transparent Antrhopic/OpenAI proxy that renders agent context (system prompt, tool docs, tool output, history) to pixels, so coding-agent CLIs like OpenCode send far fewer input tokens. Keeps tool calls and multi-turn intact. ~39% fewer tokens on real runs.
⚡ Absolute command over your AI agent's token budget — local-first compression of noisy tool output before it reaches the model. Claude Code / OpenCode / Codex / Gemini.