Smart Tree now features an intelligent, token-aware global compression system that automatically handles large outputs across the entire project. This ensures we NEVER exceed token limits while maintaining compatibility with all AI assistants.
When an MCP client connects, Smart Tree:
- Sends a small compressed test message in the initialization response
- Checks if the client acknowledges compression support
- Remembers the client's capability for the entire sessio√
// Initialization response includes:
{
"serverInfo": {
"compression_test": "COMPRESSED_V1:...",
"_compression_hint": "If you can decompress this, reply with compression:ok"
}
}The system automatically compresses when:
- Output exceeds 20,000 tokens (estimated)
- Client has confirmed compression support
- Not explicitly disabled by environment variables
Compression works everywhere:
- analyze commands (semantic, quantum-semantic, etc.)
- find operations on large codebases
- search results with many matches
- overview of massive projects
- ALL MCP tool responses
The compression manager estimates tokens using:
- 1 token ≈ 4 characters (rough estimate)
- 20,000 token threshold (keeps under 25k MCP limit)
- Automatic compression when threshold exceeded
- Format:
COMPRESSED_V1:<hex-encoded-zlib-data> - Ratio: Typically 70-90% reduction
- Use: Automatic for large outputs
- Format:
QUANTUM_BASE64:<base64-encoded-binary> - Ratio: 90-95% reduction
- Use: For quantum and quantum-semantic modes
# Disable all compression
export MCP_NO_COMPRESS=1
# Force compression always
export ST_FORCE_COMPRESS=1
# Set custom token limit (default: 20000)
export ST_MAX_TOKENS=15000# In features.toml
[compression]
max_tokens = 20000
force_compression = false
disable_compression = falseClients that decompress automatically:
- Claude Desktop (with MCP support)
- Cursor (latest versions)
- VS Code with AI extensions
- Custom MCP implementations
If client doesn't support compression:
- Smart Tree detects this automatically
- Falls back to uncompressed output
- Warns about potential token limits
- Suggests using quantum modes
# Large semantic analysis - auto-compresses if needed
analyze {mode:'semantic', path:'./huge-project'}
# Client sees compressed output only if they support it
# Otherwise, gets truncated warning# Always compress (useful for huge outputs)
analyze {mode:'semantic', compress:true}
# Or use quantum mode for maximum compression
analyze {mode:'quantum-semantic'}The compression manager tracks:
- Total compressions performed
- Bytes saved
- Estimated tokens saved
- Failed decompressions
View stats with:
st --compression-stats- ✅ Never hit token limits
- ✅ Analyze massive codebases (like Burn!)
- ✅ Get complete results, not truncated
- ✅ Automatic - no manual configuration needed
- ✅ More context in fewer tokens
- ✅ Complete project understanding
- ✅ Efficient token usage
- ✅ Automatic decompression (if supported)
- ✅ Global solution - works everywhere
- ✅ Smart detection - no breaking changes
- ✅ Token-aware - respects limits
- ✅ Statistics for optimization
- Request arrives → Check for compression acknowledgment
- Process request → Generate response
- Check response size → Estimate tokens
- Apply compression → If client supports & size exceeds limit
- Send response → With compression metadata
- Library: zlib (flate2)
- Level: Default (balanced speed/ratio)
- Encoding: Hex for text safety
- Overhead: ~50 bytes for metadata
- Client doesn't support compression
- Solution: Use
mode:'quantum-semantic'explicitly
- Client trying to display compressed data as text
- Solution: Update client or disable compression
- Very large outputs being compressed
- Solution: Use streaming or pagination
- Streaming compression - Compress chunks as they generate
- Adaptive compression - Adjust level based on content
- Client negotiation - Formal compression capability exchange
- Differential compression - Only send changes
# Before (would fail with token limit):
analyze {mode:'semantic', path:'../burn'}
# Error: MCP tool "analyze" response (44326 tokens) exceeds maximum
# After (with smart compression):
analyze {mode:'semantic', path:'../burn'}
# ✅ Auto-compressed: 177304 → 18234 bytes (89.7% reduction)
# 💡 Estimated tokens saved: 39842
# Success! Full analysis deliveredSmart Tree's global compression system ensures that:
- Token limits are NEVER exceeded when clients support compression
- Compression is automatic - no user configuration needed
- Backward compatible - non-supporting clients still work
- Global coverage - all tools benefit from compression
"Compression so smart, it knows when to squeeze!" - Aye
"Your massive codebase? We've got it covered!" - Hue