Skip to content

Latest commit

 

History

255 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SuperCompress — cut LLM context waste

SuperCompress v2

Query-aware context compression for AI applications and coding agents.
Cut coding-agent context ~64% in our B5 benchmark.

~400M Neural Keep on the hosted API · coding-agent plugin via MCP
Open weights (MIT) on Hugging Face

64.1% mean context cut on B5 (24 coding-agent cases) · 24/24 evidence passes · up to 96.6% @ ≥99% retention
(B5 coding-agent suite · evidence containment metric — not downstream LLM completion)

Get API key · Arena · Install · Benchmarks · Docs

vs Headroom · vs RTK · vs LLMLingua

PyPI npm MIT License GitHub Sponsor SuperCompress


Why it exists

Every LLM call ships a pile of context: RAG chunks, chat history, tool dumps, logs, JSON. Most of it is irrelevant to the current question — but the model still reads it, and you still pay for it.

Usual “fix” What actually happens
Truncate Deletes the middle. The answer is often in the middle.
Summarize Rewrites evidence. IDs, stack traces, and exact errors get soft.
Hope Ship the full dump. Watch the bill climb.

SuperCompress v2 compresses context against the query. It keeps answer-critical lines in their original wording and drops the rest — on our coding-agent benchmark (B5), 64.1% mean context reduction with 24/24 evidence-retention passes.


Product

SuperCompress is a compression layer in front of inference:

  1. Takes a long context + the current query
  2. Segments and scores blocks by relevance to that query
  3. Keeps entities, errors, definitions, nearby dependencies
  4. Returns a smaller prompt + token stats

The query is never compressed — only the surrounding context.

Two paths

Neural v2 (recommended) Compiler (fast/local)
Engine ~400M query-aware cross-encoder Lightweight local policy
Runtime Hosted Fly CPU (Neural Keep) Local CPU · millisecond-class
Best for Highest-quality keep on agent dumps Local preprocessing / speed
Benchmarks Launch / B5 Legacy section

Two products

Coding-agent plugin API / Python
For Cursor, Claude Code, Codex, and 40+ agent harnesses Apps, RAG, agents, backends
Install npm i -g supercompress-proxy && npx supercompress setup pip install supercompress
What you get MCP compress_context on big dumps Compress before every model call
Login Keep your normal agent login API key from the dashboard

Docs: coding agents (first 5 minutes checklist) · API quickstart

Repo map

Path What
packages/proxy Coding-agent plugin (npm)
api/ Hosted API + billing
web/ Site + docs HTML
supercompress/ Python package
docs/REPO_LAYOUT.md What belongs in OSS vs private

Private marketing, outreach, and model training stay out of this repo (see .gitignore + docs/REPO_LAYOUT.md).


Benchmarks & stats

We measure whether required evidence survives (containment), not downstream LLM completion.

Neural v2 launch (hosted ~400M) — B5 coding-agent suite

Metric Result
Mean context cut 64.1%
Evidence passes 24 / 24
Tokens 16,647 → 5,148
Max cut @ ≥99% retention 96.6%
B5 latency p50 ~5.8 s (hosted Neural Keep on Fly CPU)
Public cases (full suite) 390
Downstream LLM eval Not yet run

Aggregate mean cut across all 390 cases is ~3.6% — the engine often refuses to over-cut dense needle/QA slices. The 64.1% figure is the coding-agent suite where dumps are noisy.

Raw JSON: launch-benchmark.json · writeup: benchmarks

Legacy compiler (separate product)

CPU / millisecond-class local path. Older held-out compiler numbers (≈58–66% cut, 99.4% gold containment) live under Legacy/compiler on /benchmarks. Do not mix with Neural v2.


Try it

Coding agents (recommended)

npm install -g supercompress-proxy
npx supercompress setup

Links your account, detects agents, installs MCP. Docs: coding agents

Python / HTTP

pip install supercompress
export SUPERCOMPRESS_API_KEY=sc_live_YOUR_KEY
from supercompress.client import SuperCompress

sc = SuperCompress()
result = sc.compress(
    context=long_context,
    query="What failed and how do we fix it?",
)
print(f"{result.original_tokens} → {result.kept_tokens} tokens")
print(result.compressed_text)
curl -X POST https://www.supercompress.dev/api/v1/compress \
  -H "X-API-Key: $SUPERCOMPRESS_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"context":"...","query":"What failed?"}'

Or paste a dump into the Arena — no integration required.


How it compares

Truncate Summarize SuperCompress
Cuts tokens ✓ ✓ ✓
Uses the query ✗ weak ✓
Keeps original evidence sometimes ✗ ✓
Auditable kept lines partial ✗ ✓

More: vs truncation · vs summarization · vs alternatives


Arena   Dashboard   Agents

supercompress.dev · MIT License · Sponsor · built by Arjun Shah

About

Query-aware context compression for LLMs and coding agents. Cut ~65% of input tokens.

Resources

Contributing

Security policy

Stars

103 stars

Watchers

2 watching

Forks

Releases

Sponsor this project

Packages

Used by

Contributors

Languages