Query-aware context compression for AI applications and coding agents.
Cut coding-agent context ~64% in our B5 benchmark.
~400M Neural Keep on the hosted API · coding-agent plugin via MCP
Open weights (MIT) on Hugging Face
64.1% mean context cut on B5 (24 coding-agent cases) · 24/24 evidence passes · up to 96.6% @ ≥99% retention
(B5 coding-agent suite · evidence containment metric — not downstream LLM completion)
Get API key · Arena · Install · Benchmarks · Docs
vs Headroom · vs RTK · vs LLMLingua
Every LLM call ships a pile of context: RAG chunks, chat history, tool dumps, logs, JSON. Most of it is irrelevant to the current question — but the model still reads it, and you still pay for it.
| Usual “fix” | What actually happens |
|---|---|
| Truncate | Deletes the middle. The answer is often in the middle. |
| Summarize | Rewrites evidence. IDs, stack traces, and exact errors get soft. |
| Hope | Ship the full dump. Watch the bill climb. |
SuperCompress v2 compresses context against the query. It keeps answer-critical lines in their original wording and drops the rest — on our coding-agent benchmark (B5), 64.1% mean context reduction with 24/24 evidence-retention passes.
SuperCompress is a compression layer in front of inference:
- Takes a long context + the current query
- Segments and scores blocks by relevance to that query
- Keeps entities, errors, definitions, nearby dependencies
- Returns a smaller prompt + token stats
The query is never compressed — only the surrounding context.
| Neural v2 (recommended) | Compiler (fast/local) | |
|---|---|---|
| Engine | ~400M query-aware cross-encoder | Lightweight local policy |
| Runtime | Hosted Fly CPU (Neural Keep) | Local CPU · millisecond-class |
| Best for | Highest-quality keep on agent dumps | Local preprocessing / speed |
| Benchmarks | Launch / B5 | Legacy section |
| Coding-agent plugin | API / Python | |
|---|---|---|
| For | Cursor, Claude Code, Codex, and 40+ agent harnesses | Apps, RAG, agents, backends |
| Install | npm i -g supercompress-proxy && npx supercompress setup |
pip install supercompress |
| What you get | MCP compress_context on big dumps |
Compress before every model call |
| Login | Keep your normal agent login | API key from the dashboard |
Docs: coding agents (first 5 minutes checklist) · API quickstart
| Path | What |
|---|---|
packages/proxy |
Coding-agent plugin (npm) |
api/ |
Hosted API + billing |
web/ |
Site + docs HTML |
supercompress/ |
Python package |
docs/REPO_LAYOUT.md |
What belongs in OSS vs private |
Private marketing, outreach, and model training stay out of this repo (see .gitignore + docs/REPO_LAYOUT.md).
We measure whether required evidence survives (containment), not downstream LLM completion.
| Metric | Result |
|---|---|
| Mean context cut | 64.1% |
| Evidence passes | 24 / 24 |
| Tokens | 16,647 → 5,148 |
| Max cut @ ≥99% retention | 96.6% |
| B5 latency p50 | ~5.8 s (hosted Neural Keep on Fly CPU) |
| Public cases (full suite) | 390 |
| Downstream LLM eval | Not yet run |
Aggregate mean cut across all 390 cases is ~3.6% — the engine often refuses to over-cut dense needle/QA slices. The 64.1% figure is the coding-agent suite where dumps are noisy.
Raw JSON: launch-benchmark.json · writeup: benchmarks
CPU / millisecond-class local path. Older held-out compiler numbers (≈58–66% cut, 99.4% gold containment) live under Legacy/compiler on /benchmarks. Do not mix with Neural v2.
npm install -g supercompress-proxy
npx supercompress setupLinks your account, detects agents, installs MCP. Docs: coding agents
pip install supercompress
export SUPERCOMPRESS_API_KEY=sc_live_YOUR_KEYfrom supercompress.client import SuperCompress
sc = SuperCompress()
result = sc.compress(
context=long_context,
query="What failed and how do we fix it?",
)
print(f"{result.original_tokens} → {result.kept_tokens} tokens")
print(result.compressed_text)curl -X POST https://www.supercompress.dev/api/v1/compress \
-H "X-API-Key: $SUPERCOMPRESS_API_KEY" \
-H "Content-Type: application/json" \
-d '{"context":"...","query":"What failed?"}'Or paste a dump into the Arena — no integration required.
| Truncate | Summarize | SuperCompress | |
|---|---|---|---|
| Cuts tokens | ✓ | ✓ | ✓ |
| Uses the query | ✗ | weak | ✓ |
| Keeps original evidence | sometimes | ✗ | ✓ |
| Auditable kept lines | partial | ✗ | ✓ |
More: vs truncation · vs summarization · vs alternatives
supercompress.dev · MIT License · Sponsor · built by Arjun Shah