Compress outputs and context before LLMs to cut tokens without losing answers.
Copy the install command and let the AI configure it · recommended for beginners
No copy-paste install info for "headroom" yet — see the docs or source repo.
Compress these RAG retrieval chunks to 20% of their original length without losing key facts, numbers, or citation relationships, and output streamlined context suitable for sending to an LLM.
A highly compressed retrieval context that preserves essential information and reduces downstream prompt token usage.
Compress the following agent execution logs, tool outputs, and intermediate steps into the smallest useful summary, preserving errors, key parameters, final results, and dependencies for continued LLM reasoning.
A concise log summary that keeps the critical details needed for troubleshooting and reasoning.
Read this large document or code file, then compress it aggressively while preserving structure, key conclusions, and important snippets, and output a version optimized for LLM consumption.
A shorter but high-density representation of the file that an LLM can process at lower cost.
Compress prompts, tool outputs, and replies to reduce LLM token costs.
Compress MCP tool schemas to cut tokens while preserving semantics deterministically.
Compress long contexts and retrieve reusable summaries to reduce LLM token usage.
Compresses LLM conversation context while preserving meaning and reducing token usage.
Optimize Claude Code token usage to cut costs and improve development efficiency.
Cut AI API costs dramatically with token measurement, compression, caching, and pruning.