Context window compression and management utilities
-
Updated
Feb 11, 2026 - Python
Context window compression and management utilities
Ultra-compressed language protocol for LLMs — no retraining needed. Reduce token usage by 8x-10x.
See what's eating your LLM context window — analyze prompts and MCP tool manifests for token cost, per-tool breakdown, and what to lazy-load. Zero dependencies.
Measured context-engineering notes for local-model agent builders — context diets, re-reads, token math — adapted from KyaniteLabs/context-kit with every number's validity label intact. The canonical repo stays the instrument; these notes restate its measured results for harness builders.
Agentic Workflow Based Coder System
Measure what your Claude Code setup costs before any work happens. Fixed per-turn overhead with @imports followed, and model-pin drift across all eight precedence layers. Reads and prints only.
OpenCode TUI plugin: counts the turns cut off mid-thought, says whether the output cap or the context window did it, and shows what thinking costs per turn and after which tool.
Streamline long GPT-5.5 sessions for researchers and developers — instantly compress 1M-token context into focused summaries for continuity.
See where your Claude Code tokens go — and stop the worst of it. A Claude Code plugin.
What your coding agent's context window is made of — and what it actually costs. Claude Code, Codex, OpenCode, Devin.
Surge an AI endpoint: does an oversized prompt error cleanly or silently truncate? Plus load under concurrency — latency percentiles, error rate, capacity knee.
Measure what your AI agent loads before you type a word: the context floor. Per-file attribution across Claude Code/Codex/Hermes/OpenClaw, per-tokenizer honesty, fixtures-as-contract, CI budget gate.
Prompt and RAG context compression with a fidelity evaluator — 48% fewer tokens at 100% fact recall, where naive truncation scores 21.7%. Zero dependencies.
Lampka kontrolna lokalnego modelu: co robi Ollama i czy nie ucięła Twojego promptu. A one-screen live view of local Ollama — GPU, memory, and silent prompt truncation.
AI Agent Memory: Generate & optimize AGENTS.md, prevent context bloat, and configure memory-cache.
✂️ Lightweight, framework-agnostic Python library for intelligent LLM context window management and memory heap packing. Zero dependencies.
Checkpoint and rehydrate for AI agents: keep the decisions, paths, and dead ends a context-limit summary drops.
Fleet observability for Claude Code — one table for every session you have running. Reads documented surfaces only, never an internal format.
Message 50 costs 6x message 1. A Claude Code statusline that makes the snowball visible.
To associate your repository with the context-window topic, visit your repo's landing page and select "manage topics."