JuliusBrussee/cavemanPublic

🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.

AI summary: An agent skill and local proxy that reduces LLM token consumption by condensing context and output into concise 'caveman' speech.

Stars
109.8K
+272 today
Forks
6.3K
Watchers
253
Open issues
86
Open PRs
81
Contributors
~61
Commits
823
Branches
45

GoApache-2.0Created Apr 4, 2026Last push todayLatest release v3.0.0+1.7K stars this week+6.2K this month

Quick answers

What is caveman?
An agent skill and local proxy that reduces LLM token consumption by condensing context and output into concise 'caveman' speech.
What does caveman do?
Caveman is a prompt engineering tool designed to drastically reduce the number of tokens consumed by coding agents like Claude Code, Cursor, and Codex. It operates through two distinct components: a local proxy that actively shrinks the context (like tool schemas, files, and history) sent to the provider, and an agent skill that forces the LLM to output highly condensed, 'caveman-like' responses while preserving byte-exact code and commands. By eliminating conversational filler and aggressively truncating context, it achieves a measured 33.2% reduction in provider-reported input tokens. This approach maintains the agent's core reasoning capabilities and factual accuracy while significantly lowering API costs and latency.
Who is caveman for?
Power users of AI coding assistants and developers managing high API costs. It is targeted at users who prefer direct, unembellished answers and want to optimize their token usage across tools like Claude Code or Cursor.
How do I get started with caveman?
npm install -g @caveman-ai/cli && caveman setup --install
How popular is caveman on GitHub?
JuliusBrussee/caveman has 109,775 stars and 6,348 forks on GitHub, and gained 1,708 stars in the last 7 days.
What license does caveman use?
JuliusBrussee/caveman is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
050K100KJul 2026Aug 2026Sep 2026Oct 2026
109.8K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 6 commits2026-04-05: 10 commits2026-04-06: 9 commits2026-04-07: 9 commits2026-04-08: 16 commits2026-04-09: 18 commits2026-04-10: 3 commits2026-04-11: 16 commits2026-04-12: 7 commits2026-04-13: 1 commit2026-04-14: 6 commits2026-04-15: 5 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 1 commit2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 1 commit2026-05-01: 12 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 17 commits2026-05-11: 0 commits2026-05-12: 3 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 1 commit2026-05-19: 0 commits2026-05-20: 2 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 9 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 1 commit2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 1 commit2026-06-11: 0 commits2026-06-12: 7 commits2026-06-13: 0 commits2026-06-14: 5 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 1 commit2026-07-01: 8 commits2026-07-02: 14 commits2026-07-03: 4 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 6 commits2026-07-22: 0 commits2026-07-23: 1 commit2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 1 commit2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 1 commit2026-08-02: 1 commit2026-08-03: 4 commits2026-08-04: 1 commit2026-08-05: 2 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 4 commits2026-08-09: 1 commit2026-08-10: 1 commit2026-08-11: 6 commits2026-08-12: 6 commits2026-08-13: 7 commits2026-08-14: 1 commit2026-08-15: 9 commits2026-08-16: 8 commits2026-08-17: 1 commit2026-08-18: 10 commits2026-08-19: 39 commits2026-08-20: 16 commits2026-08-21: 6 commits2026-08-22: 0 commits2026-08-23: 27 commits2026-08-24: 1 commit2026-08-25: 11 commits2026-08-26: 1 commit2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 13 commits2026-08-30: 53 commits2026-08-31: 0 commits2026-09-01: 2 commits2026-09-02: 14 commits2026-09-03: 22 commits2026-09-04: 5 commits2026-09-05: 4 commits2026-09-06: 1 commit2026-09-07: 5 commits2026-09-08: 41 commits2026-09-09: 2 commits2026-09-10: 4 commits2026-09-11: 3 commits2026-09-12: 1 commit2026-09-13: 6 commits2026-09-14: 44 commits2026-09-15: 9 commits2026-09-16: 10 commits2026-09-17: 4 commits2026-09-18: 5 commits2026-09-19: 4 commits2026-09-20: 4 commits2026-09-21: 2 commits2026-09-22: 0 commits2026-09-23: 2 commits2026-09-24: 21 commits2026-09-25: 8 commits2026-09-26: 0 commits2026-09-27: 35 commits2026-09-28: 23 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 3 commits2026-10-02: 0 commits2026-10-03: 0 commits
715 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Landmark project

    109,775 stars

  • Very active

    715 commits in 52 weeks

  • Well documented

    High community health score

  • Permissive license

    Apache-2.0

  • Repeat trending

    39 trending appearances

  • Top 10% tracked

    Rank 62 of 1135

What caveman does

Caveman is a prompt engineering tool designed to drastically reduce the number of tokens consumed by coding agents like Claude Code, Cursor, and Codex. It operates through two distinct components: a local proxy that actively shrinks the context (like tool schemas, files, and history) sent to the provider, and an agent skill that forces the LLM to output highly condensed, 'caveman-like' responses while preserving byte-exact code and commands. By eliminating conversational filler and aggressively truncating context, it achieves a measured 33.2% reduction in provider-reported input tokens. This approach maintains the agent's core reasoning capabilities and factual accuracy while significantly lowering API costs and latency.

Power users of AI coding assistants and developers managing high API costs. It is targeted at users who prefer direct, unembellished answers and want to optimize their token usage across tools like Claude Code or Cursor.

  • Context Shrinking Proxy: Intercepts and compresses tool schemas, files, and conversation history before they reach the provider API.
  • Concise Output Generation: Forces the agent to eliminate conversational boilerplate and answer using extreme brevity.
  • Byte-Exact Code Preservation: Ensures that while natural language is compressed, all generated code, terminal commands, and system errors remain perfectly intact.
  • Broad Agent Compatibility: Works natively across more than 30 different coding agents, including Claude Code, Gemini CLI, and OpenCode.
  • Flexible Deployment: Offers both a local proxy for input compression and a direct agent skill for output reduction, which can be used independently or together.

Where teams use it

Cost Reduction for AI Coding

Helps developers minimize API billing costs by stripping unnecessary tokens from both the input context and the model's responses during continuous coding sessions.

Latency Optimization

Accelerates agent response times by reducing the total volume of text the LLM needs to process and generate, focusing compute entirely on the actionable code.

Streamlined Terminal Output

Cleans up the developer's terminal interface by preventing the agent from generating long, conversational explanations for simple code fixes.

Maximized Context Windows

Allows developers to fit larger codebases or more extensive chat histories into a model's context window by aggressively compressing the non-essential text.

Getting started: npm install -g @caveman-ai/cli && caveman setup --install

README

main branch
Caveman

why use many token when few do trick

Your AI coding agent bills by the word and writes like it knows that. Caveman make it stop.

ThePrimeagen reacts to Caveman: No way this actually works

▶️ ThePrimeagen reacts: "No way this actually works"

GitHub stars npm downloads middleware on npm middleware on PyPI 30+ agents 10 native wrap profiles License skills.sh

🏆 #1 on GitHub Trending · July 2026  ·  🥇 #1 Repository of the Day on Trendshift · April 2026

#1 on Hacker News  ·  #8 Product of the Day

📄 Cited in CAVEWOMAN, an Adobe Research paper that measured caveman-style output cutting cost 1.4 to 2.4×, up to 3×  ·  🧪 Tested by JetBrains on 86 real coding tasks: "costs you nothing measurable in quality"

Caveman - why use many token when few do trick | Product Hunt JuliusBrussee%2Fcaveman | Trendshift

⚡ One command, no account, no API key. npx skills add JuliusBrussee/caveman -g → Quick Start



🪨 See it

🗣️ Normal agent · 69 tokens Caveman agent · 19 tokens

The reason your React component is re-rendering is likely because you're creating a new object reference on each render cycle. When you pass an inline object as a prop, React's shallow comparison sees it as a different object every time, which triggers a re-render. I'd recommend using useMemo to memoize the object.

New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.

Same diagnosis. Same fix. Same useMemo. The only thing that died was the throat-clearing.

Code, commands, file paths, and exact error messages never get cavemanned. Only the prose around them does. Security warnings and "are you sure?" confirmations come back in full sentences on their own, then caveman resumes.

Caveman no make brain smaller. Caveman make mouth smaller.

Half the fun is that your agent talks like it just discovered fire. The other half is that it is still right.


🌍 Why this exists

A token is what AI billing counts, roughly three quarters of a word. Your agent pays for every token it writes and every token it reads. Most agents write like a cover letter and read like a firehose.

Caveman attacks both ends, in the agent you run and in the one you build:

  • The skill shrinks what the agent says. One rule file. Free forever. Works in 30+ agents.
  • The proxy shrinks what the agent reads: logs, test output, JSON, diffs, search results. Runs on your machine. Every squeezed byte gets a backup, so the agent can always pull the original back.
  • The middleware does the same inside your own code: one wrapper around the LangChain, Vercel AI SDK, OpenAI, or Anthropic call you already make. Tool results get shrunk before the model sees them, the original stays in your history, and the model can fetch it back.

Started as a joke on a Friday in April 2026. Hit 4,000 stars in a week. Now past 100,000, with a research paper, a JetBrains lab test, and a Primeagen reaction video. The joke got serious. The voice did not.


⚡ Quick Start

Caveman come in two sizes. Start small.

Small rock: the skill

A rule file that makes your agent answer in caveman. Apache-2.0, free forever, works in 30+ agents (Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot, more). One command:

npx skills add JuliusBrussee/caveman -g

Type /caveman if your agent doesn't wake up on its own. That the whole install. One rock.

Big rock: the proxy

Runs on your machine, between your agent and the AI provider, and shrinks what the agent reads before every call. Apache-2.0, CLI and runtime both:

npm install -g @caveman-ai/cli && caveman setup --install
caveman claude        # or codex · gemini · aider · kilo · qwen · opencode · hermes · openclaw · pi

Your own app: the middleware

Building an agent in code instead of running one in a terminal? Same shrinking, one wrapper around the call you already make. Apache-2.0 client, stable 1.0:

npm install @caveman-ai/middleware @caveman-ai/sdk        # TypeScript, plus your framework (ai, openai, …)
pip install 'caveman-middleware[langchain]' caveman-sdk   # Python 3.11+, swap the extra for your framework

Six lines of code and a local runtime. Full walkthrough below.

They stack. Most people start with the small rock and graduate.

More doors into the cave · full installer, Windows, single agents, uninstall

The full installer wires up Claude Code hooks and the statusline badge, finds every supported agent on your machine, and skips agents you no have. Safe to re-run. Needs Node.js 22.13+.

curl -fsSL https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.0.0/install.sh | bash

Windows, PowerShell 5.1+:

irm https://raw.githubusercontent.com/JuliusBrussee/caveman/v3.0.0/install.ps1 | iex

Just one agent:

# Claude Code
claude plugin marketplace add JuliusBrussee/caveman && claude plugin install caveman@caveman

# Gemini CLI
gemini extensions install https://github.com/JuliusBrussee/caveman

# Oh My Pi (OMP)
npx -y github:JuliusBrussee/caveman -- --only omp

# Qwen Code CLI, then its Caveman wrapper
npm i -g @qwen-code/qwen-code
caveman qwen

# Codex, Cursor, Windsurf, Cline, and other skills-compatible agents
npx skills add JuliusBrussee/caveman --skill '*' -a codex --yes -g  # replace codex with your agent profile

Install broke? Open your agent in this repo and say: "Read CLAUDE.md and INSTALL.md, install caveman for me." Agent read repo, agent fix own brain. Snake eat tail.

Changed your mind: npx -y github:JuliusBrussee/caveman -- --uninstall

The full 30+ agent matrix, dry runs, flags, and verification live in INSTALL.md.

🕐 The first five minutes

Small rock. The skill, right after npx skills add:

  1. Ask it something. Any coding question. Watch the preamble vanish and the answer stay.
  2. Turn the dial. /caveman lite for tight-but-polite. /caveman ultra for grunts. /caveman wenyan for classical Chinese, because someone asked.
  3. Commit like a caveman. /caveman-commit writes a Conventional Commit in one line.
  4. Review like a caveman. /caveman-review gives one finding per line: L42: 🔴 null deref. Guard it.
  5. Shrink your memory files. /caveman-compress CLAUDE.md cuts the prose, keeps every heading, path, and command, and backs up the original.
  6. Come home. Say stop caveman. Normal prose returns. No hard feelings.

Big rock. The proxy, right after npm install -g @caveman-ai/cli:

  1. Find out where your tokens go. caveman learn reads months of agent history already on your disk, locally, and ranks the places your tokens go, biggest first, with a one-line fix behind each. Do this before anything else. It is the most useful five minutes in this README. After step 3 it keeps watching by itself and speaks up only when something new appears.
  2. Let it fix them. caveman learn implement hands each fix to Claude Code or Codex one diff at a time, applied only on your yes, and undoes anything that did not make each message smaller.
  3. Wrap your agent. caveman claude (or codex, gemini, aider, opencode, pi, …) puts the proxy in front of it. Logs, test output, JSON, and diffs get shrunk before the provider sees them. Originals stay on disk, and the agent can pull any of them back.
  4. Shrink the noisy stuff. caveman shrink -- pnpm test compresses command output. caveman browse <url> gives the agent a compressed view of a web page instead of a 15,000-token accessibility dump.
  5. Prove it on your own work. caveman trial -- claude runs a real session with and without caveman, then caveman trial report shows the difference. That A/B outranks every number on this page. A trial needs its own proxy, so if you already did step 3 it will tell you to run caveman disable claude first, and caveman enable claude after. Caveman rather say "cannot measure this" than hand you a report full of zeros.
  6. Shrink caveman itself. caveman convert --dry-run shows which installed skills get cheaper as PNG pages the model reads as an image. Convert the profitable ones, revert byte-for-byte any time.
  7. Watch the bill. caveman stats for history and estimates. /caveman-stats inside Claude Code for that session.

📊 The Numbers

Every number below is either from a committed run in this repo or from a named third party. Nothing rounded up. Where a number is small, it says so. Where a row is red, it stays red.

What the skill saves (writing less)

Who measured What they measured Result
Adobe Research (CAVEWOMAN, arXiv 2606.24083) Eight models, five datasets, five compression levels Output-side caveman style cuts realized cost 1.4 to 2.4× per model, up to 3× in the best case
JetBrains 86 real coding tasks, paired A/B, Claude Code 2.1.200. Skill only, no proxy (July 2026, before the proxy existed) 8.5% fewer output tokens, about 10% cost. No detectable quality change (sign test p = 0.82)
This repo (committed eval snapshot) Ten dev questions, skill vs a plain Answer concisely. control, claude-opus-4-6 50% fewer output tokens at the median on top of the terse control. Length only, not correctness

Read those three together and you get the honest picture. Chat-style Q&A: big cut. Agentic coding sessions, where most tokens are code and tool calls that the skill never touches: high single digits on output, quality flat.

The JetBrains number is why the proxy exists. They measured the skill alone, in July 2026, before the proxy shipped. Their finding was that an agent's bill is mostly reading, not writing, and no talking style fixes that. So we built the thing that shrinks the reading. The table below is what that changed.

The Adobe paper's other finding matters too: compressing the human's prompt into caveman-speak makes models answer longer and worse. Caveman never rewrites your prompts. Only the agent's mouth.

The rules add input tokens on every call, and whether shorter output pays for them depends on your agent, caching, and billing. Full accounting: docs/HONEST-NUMBERS.md.

No reviewed API benchmark result is published here yet. Run uv run python benchmarks/run.py to generate a new result, then review its raw response pairs and quality before publishing the generated table.

What the proxy saves (reading less)

Your agent rereads logs, test output, diffs, and half your repo all day. The proxy shrinks that stream before it reaches the provider. Pinned 54-run Claude Code benchmark, provider-reported input tokens, three runs per case, every answer checked against an exact oracle:

Case Direct Claude Code Through caveman Change
CSV outlier hunt 165,823 74,484 -55.1%
Log needle in haystack 148,807 74,068 -50.2%
YAML config drift 132,124 71,027 -46.2%
Test output failure 150,377 108,514 -27.8%
Deployment JSON drift 147,975 108,939 -26.4%
Dashboard HTML alert 140,687 154,641 +9.9%
Total 885,793 591,673 -33.2%

18 of 18 answer checks passed. Case-clustered 95% interval: 14.6% to 48.5%. In the same suite, Headroom's wrap saved 6.7% and failed 3 of 18 checks. Method, provenance hashes, and limits: docs/WRAP-BENCHMARK.md. Raw harness artifacts are not in this checkout, so treat it as a pinned report, not a public reproduction.

Maintainer note. The HTML row is red and it stays red. That case had no compression transform, so caveman paid its own overhead and won nothing back. The day I hide a red row is the day you should stop trusting the green ones.

Everything else caveman shrinks

Surface Measured Number
Browser pages Focused question against a 200-row table, vs the Playwright ARIA snapshot 121 tokens vs 15,704. 129.8× smaller. Tiny forms lose 2.3×; the benchmark says so
Memory files (/caveman-compress) Five real CLAUDE.md-style fixtures 46% smaller on average, headings, code, paths, and URLs verified intact
The skill itself (pixel mode) Rendered to PNG pages the model reads as an image 1,069 to 415 estimated tokens, a 61% cut
Your harness prefix (subagent-tax) What every subagent re-sends before doing any work On one real machine, 219k of a 267k-char request was tool schemas. Run it on yours

🧮 How it compares

Many tool in valley promise small token. They work at different layers, so first what each one touches, then what got measured. Every quote below is from that tool's own README or GitHub page on 2026-09-19.

What each one touches

Tool What it shrinks Get the original back? Phones home
Caveman What the agent says (skill) and what it reads: tool output, logs, JSON, diffs, test output, web pages (proxy) Always. Byte-exact original in local SQLite, one recovery handle CLI and its agent hooks: usage stats with a random install ID and your IP, on by default, caveman telemetry off. Skill alone: never
RTK Shell command output only: ls, cat, grep, git, test runners. Read and Grep tool calls bypass it When a command fails or gets cut short, or opt-in for successful runs Off by default, opt-in
Headroom Tool output, logs, files, and history, through a local proxy Yes, reversible cache On by default, HEADROOM_BEACON=off
context-mode Tool output, run in a sandbox so raw data never enters context Matching sections from a searchable index, not the whole thing back Never
pxpipe Text context, re-rendered as images the model reads No. "It is lossy." Misses are silent Local log only

What the tin says, and who checked

Tool Says on the tin Who checked, on what Found
Caveman Only what this page measures This repo, pinned 54-run Claude Code suite, every answer checked against a known-right answer 33.2% fewer input tokens, 18/18 answers right
Caveman, skill only JetBrains, 86 real coding tasks, paired A/B 8.5% fewer output tokens, quality flat (sign test p = 0.82)
RTK "cuts up to 90% of the bash output your agent reads". Their README adds: "it is not the same as cutting your bill by 90%" JetBrains, same lab, same method, 86 tasks, 425 billed trials +7.6% median cost per task at low reasoning effort (p = 0.004), +0.1% at high. Quality tie
Headroom "20% fewer tokens for coding agents, 60-95% fewer tokens for JSON" This repo, same 54-run suite as above 6.7% fewer input tokens, 15/18 answers right
context-mode "315 KB becomes 5.4 KB. 98% reduction." Own size numbers only. No quality check published —
pxpipe "~59–70% lower end-to-end bill" Own SWE-bench runs Lite 10/10 both arms. Pro 14/19 with, 15/19 without, and their rerun of the one split says run-to-run variance

Same suite, same model, same questions

The one place two of these tools ran side by side against the same known-right answers. Claude Code 2.1.223, claude-sonnet-5, Headroom 0.33.0, six agent-shaped workloads, three runs each, provider-reported input tokens:

Arm Answers right Provider input tokens vs direct 95% interval
Direct Claude Code 18/18 885,793 baseline
Caveman wrap + skill 18/18 591,673 -33.2% 14.6% to 48.5%
Headroom wrap 15/18 703,202 on its 15 correct runs -6.7% on those 15 -0.7% to 17.9%

Caveman used fewer tokens in 15 of the 18 paired runs. Headroom's 703,202 covers only the 15 runs it answered right, so its 6.7% is against those same 15 direct runs, not against the 885,793 total. Its three failed YAML runs stay in the table and count for nothing. Caveman's one red row, HTML at +9.9%, is in the per-case table above and stays red too. We ran this ourselves, and the raw harness artifacts are not published yet, so it is a pinned report, not something you can re-run from this repo. Method and hashes: docs/WRAP-BENCHMARK.md.

RTK, context-mode, and pxpipe were not in that run. RTK rewrites shell output, and this suite hands the agent its data through a tool call, not the shell, so RTK would have sat idle. Different layer, different test. Fair is fair on the rest: RTK's telemetry is opt-in and ours is opt-out, context-mode sends nothing anywhere, and pxpipe ran SWE-bench where we have not. Stack them if you like. Headroom's own README lists caveman as something it happily runs behind.


📣 In the Wild

ThePrimeagen · "No way this actually works"
Full reaction on The PrimeTime →

Adobe Research · CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
Adeyemi, Rossi, Dernoncourt · arXiv, June 2026 · cites this repo. The style is now a benchmarked register.

JetBrains · Speaking to AI Agents like Cavemen Saves 65% of Tokens. We Test.
The most rigorous outside A/B so far, run on the skill alone before the proxy existed. Their verdict: "Use it if you like it. It is fun, and it costs you nothing measurable in quality." Their 8.5% is the number that made us build the proxy.

Hacker News · #1, 904 points, 366 comments

The New Stack · Getting Claude Code to grunt in Caveman-speak might not save as many tokens as you think
Fair headline. We link it anyway. See The Numbers.

GitHub Trending · #1 overall, July 2026
Trendshift · #1 Repository of the Day (April 2026) · #1 JavaScript repo of the month (April) · #1 Go repo of the month (July)

Product Hunt · #8 Product of the Day

Star History Chart


💬 The skill, unpacked

One rule file, one talking style, plus a small toolbox. /caveman lite|full|ultra|wenyan-lite|wenyan-full|wenyan-ultra sets intensity. /caveman off or normal mode turns it off.

Level Same question: "Why does my React component re-render?"
lite Your component re-renders because you create a new object reference each render. Wrap it in useMemo.
full (default) New object ref each render. Inline object prop = new ref = re-render. Wrap in useMemo.
ultra Inline obj prop, new ref, re-render. useMemo.
wenyan-full 每繪新生對象參照,故重繪;以 useMemo 包之則免。

Three things the skill will never do: shorten your code, paraphrase an error message, or grunt through a security warning. It drops to full sentences for anything irreversible, then picks the club back up.

Everything in the box · commit messages, reviews, subagents, work patterns
Tool / command What you get
/caveman [lite|full|ultra|wenyan-lite|wenyan-full|wenyan-ultra|off] Shorter replies at the intensity you choose.
cavecrew-investigator, cavecrew-builder, cavecrew-reviewer Compressed subagent presets for locating, editing, and reviewing code.
/caveman-commit Terse Conventional Commit messages.
/caveman-review One-line, actionable review findings.
/caveman-compress <file> Smaller Markdown memory files, with the original backed up.
/caveman-stats Recorded Claude Code token usage; savings unknown without a measured comparison.
/caveman-help One-screen reminder of every mode and command.
investigate-first, lean-build, surgical-patch, safe-refactor, migration, verify-and-stop Work patterns that write less code, so the agent bills fewer tokens. Your agent picks these up on its own when a task fits.
/caveman-setup, /caveman-discover, /caveman-learn, /caveman-manage, /caveman-optimize, /caveman-explore, /caveman-evidence-review Drive the caveman engine and proxy: set it up, find where tokens go, act on what it finds.

🔧 The proxy, unpacked

One local process. Your agent talks to it, it talks to your provider. No Caveman server in the path, and your Claude Pro/Max login passes through to Anthropic untouched. Originals of everything it compresses sit in a SQLite file on your machine with a recovery handle, so the agent can always ask for the full version back.

 Your agent  (Claude Code · Codex · Gemini · Aider · opencode · Pi · …)
      │   tool output · logs · JSON · diffs · search results
      ▼
 ┌────────────────────────────────────────────────────┐
 │  caveman proxy   (your machine, your keys)          │
 │  detect() → json · log · code · diff · search · text│
 │  originals → local SQLite, recovery handle returned │
 └────────────────────────────────────────────────────┘
      │   smaller prompt, same answer
      ▼
 Your provider  (Anthropic · OpenAI · Google · Bedrock · Vertex · Azure · OpenRouter)

Whole team? One container. Same proxy in your VPC, one shared token, keys stay on server. Deploy it →

Terminal demo: caveman compress reads a large JSON payload and emits a much smaller compressed version, byte-exact recoverable

What the engine keeps, by payload type · and the wrap stack diagram

coding agent talks to a local caveman proxy that forwards upstream to the provider with auth passed through byte-exact; a CCR store below the proxy keeps the original bytes and returns a recovery handle to the agent; an MCP toolkit side-channel gives the agent caveman_retrieve, toon encode/decode, and browse

detect() types each payload and routes it to a compressor that keeps what answers depend on:

Detected type Keeps Target savings
json keys, structure, error/message subtrees; collapses repetitive arrays 70-90%
log errors, stack traces, first/last lines; drops INFO and progress noise 85-95%
code imports, signatures, types; elides function bodies, syntax stays valid 40-70%
diff file/hunk headers and changed lines; elides repeated context 60-80%
search-result top/bottom hits plus diagnostic/security hits 80-95%
text / HTML headings, opening/closing context, important sections 50-80%

contextwindow.Pack() additionally fits candidate context into a token budget by BM25 relevance, recency, and error signal, returned in original order so chronology survives.

Any MCP host gets the same powers through five tools: caveman_compress, caveman_retrieve, caveman_stats, caveman_toon_encode, caveman_toon_decode.

Where your tokens go

Months of your agent history already sit on your disk. caveman learn reads it, locally, read-only, no account, and ranks the places your tokens go, biggest first, with a one-line fix behind each. Then it keep watching, so you not have to remember.

Your agent re-sends its instructions (CLAUDE.md, skills, hooks) and the whole conversation with every message it sends the model. So a few hundred extra tokens in a setup file get paid again on every single message. That is what learn hunts.

caveman learn             # Claude Code + Codex + Gemini CLI + opencode; aider via CAVEMAN_AIDER_ROOT
caveman learn implement   # hand the fixes to Claude Code or Codex, one diff at a time, applied only on your yes

Caveman Learn report: a short summary and savings cards on the left; the biggest places tokens go, with one fix opened, and a chart of how full sessions get on the right

It run itself. Once caveman is on your agent (caveman claude, caveman codex, …), learn re-scans quietly after a session ends. Low priority, at most every 6 hours, never makes the session wait. When something new and heavy shows up, your next session opens with one line, one time:

caveman learn: new finding — Project CLAUDE.md is 423 lines (~9,699 tokens), loaded with every message (~9.7k tokens in every message, estimate). Run `caveman learn` to review.

Nothing new, nothing said. caveman learn autopilot off if you rather run it by hand.

It check your memory files. Claude Code loads only the first 200 lines (or 25KB) of MEMORY.md. Everything past that, your agent never sees, and nothing tells you. Learn tells you. It also catches @imports pointing at files that are gone, the same rule pasted into two files your agent loads (you pay for it twice, every message), file paths in CLAUDE.md that no longer exist, and memory notes the index forgot to link.

It say if you getting better. Week over week, from your own sessions, first run included. Real output from the maintainer's machine, bad news left in:

last 6 weeks  tokens per session   ▄▂▃▁█┊▄  +185% · worse
              peak context used    ▁▁▄▅█┊▁  +4 points · worse
              overloaded messages  ▁▂▃▇█┊▅  +2.9 points · worse
              week of Sep 21 (928 sessions) vs the 4 weeks before
              a trend is not a saving, and it does not show the cause

"Overloaded" means the conversation filled more than half of what the model can hold at once. Past that, answers tend to get worse (a common rule of thumb). Each week counts its middle session, not the average, so one giant session can't skew it. Weeks under 5 sessions say "not enough data" instead of guessing. The ┊ marks the week still running: shown, never compared.

It prove the fix, or undo it. implement re-measures after every change and undoes anything that didn't make each message smaller. Some fixes can't be re-counted, like a new skill that only pays off when it gets used. For those, caveman learn experiment runs it on for a stretch and off for a stretch over your own sessions, and gives no verdict before 5 sessions each way. Caveman never makes your agent dumber to make it cheaper.

Every verb, every check, every number it will and won't show: docs/technical/learn.md.

More verbs

caveman explore install         # read-only FastContext subagent: finds code as path:line
caveman shrink -- pnpm test     # compress noisy command output, byte-exact recoverable
caveman browse <url>            # local Chrome over a compressed a11y tree
caveman mem remember|recall     # durable memory; `mem recover <handle>` = original bytes
caveman trial -- claude         # A/B a real session, then `trial report` (needs `disable` first)
caveman toon encode|decode      # the TOON re-encoder, standalone
caveman stats                   # token history, API estimates, subscription equivalents

Pixel mode

Caveman eating its own tail. Every skill you install is prompt text your agent reloads on every call. caveman convert renders the skill body to PNG pages in place, and the model reads it as an image. On the caveman skill itself: 1,069 to 415 estimated tokens, a 61% cut.

caveman convert --dry-run        # every installed skill, with the token math, no writes
caveman convert --agent claude   # convert the profitable ones
caveman convert --revert         # byte-identical restore from SKILL.orig.md

Convert only fires when pages beat the text. Any failure leaves the skill byte-identical and names the gate that said no.

Wrap any agent

caveman <agent> turns the proxy on for good and launches the agent. caveman wrap <agent> runs one session and leaves nothing behind. It never edits your config files.

In managed mode the Claude Code wrap also sends the repository (github.com owner/name) and current branch name as x-cave-tags, so Cloud can join a session's spend to the change it shipped. Branch names can carry a person's or customer's name. Not OK? Set CAVEMAN_WORK_TAGS=0 and no tags go. Or set your own x-cave-tags in ANTHROPIC_CUSTOM_HEADERS: the wrap sends yours exactly as written and adds nothing.

Agent Vendor How it's wrapped
Claude Code Anthropic env vars
OpenAI Codex CLI OpenAI env vars (API key) · ephemeral CODEX_HOME (ChatGPT login)
Gemini CLI Google env vars
Aider OpenAI/Anthropic env vars
Kilo Code Kilo Code KILO_CONFIG_CONTENT, your kilo.json untouched
Qwen Code QwenLM ephemeral system-settings overlay, source settings untouched
opencode sst inline config via env, your opencode.json untouched
Hermes Agent Nous Research --provider custom + env
OpenClaw OpenClaw ephemeral merged config, your config read-only
Pi pi.dev bundled native extension, your ~/.pi config untouched
Fine print · tested versions, default loadout, SDK recipes

Tested against real sessions on Hermes v0.18.0, OpenClaw 2026.6.11, Pi 0.84.2, Kilo Code 7.5.6 (the CLI, not the editor extension), and Qwen Code 0.22.3. Persistent shortcuts are journaled and reversible with caveman disable <agent>.

OpenClaw, for the record, is a lobster. Lobster claw still sharp. Lobster mouth now small.

The default wrap hands the agent the five MCP tools, the browse server when Chrome resolves, command-output shrink on Claude, opencode, Gemini, Hermes, and OpenClaw, and pixel mode on new skill installs. Codex skips the shrink hook because its runtime rejects the rewrite (openai/codex#18491). Turn pieces off in ~/.caveman-cloud/config.json.

Agent not on the list, or building your own? Wrap one call natively with the middleware, or point any provider SDK or framework (Vercel AI SDK, LangChain, LiteLLM, OpenAI Agents, CrewAI, PydanticAI) at the local proxy with a baseURL swap: integrations/recipes/. New native agent is usually one JSON profile in agents/profiles/.


🧩 Caveman in your own app

The proxy shrinks what a coding agent reads. The middleware does the same thing for the agent you are building, in the framework you already use. One wrapper around one call. Before each provider request it swaps big tool results for a shorter copy and hands the model a caveman_retrieve tool, so the model can read the original back whenever the short copy is not enough. Your conversation history keeps every original byte. Your provider, your client, your retries, your streaming: untouched.

TypeScript (Vercel AI SDK shown):

+import { createMiddlewareRuntime } from "@caveman-ai/sdk/middleware";
+import { withCaveman } from "@caveman-ai/middleware/ai-sdk";
+
+const runtime = createMiddlewareRuntime({ endpoint: "http://127.0.0.1:8787", mode: "compress" });
+const scope = { namespace: "support", session_id: conversationId, branch_id: "main", cache_epoch: "0" };
+
-const result = streamText(options);
+const result = streamText(withCaveman(options, { runtime, scope }));

Python (LangChain shown):

+from caveman_cloud.middleware import MiddlewareRuntime, Scope
+from caveman_middleware.langchain import with_caveman_agent
+
+runtime = MiddlewareRuntime(endpoint="http://127.0.0.1:8787", mode="compress")
+scope = Scope("support", conversation_id, "main", "0")
+
-agent = create_agent(model=model, tools=tools)
+agent = create_agent(**with_caveman_agent({"model": model, "tools": tools}, runtime=runtime, scope=scope))

The runtime is the same local proxy from the big rock, started once beside your app:

npm install -g @caveman-ai/cli && caveman setup --install
CAVEMAN_MODE=compress caveman start     # binds 127.0.0.1:8787; plain `caveman start` only records
Frameworks with a native adapter
TypeScript @caveman-ai/middleware Vercel AI SDK · OpenAI · Anthropic · Google GenAI · LangChain · Strands · Mastra · MCP
Python caveman-middleware OpenAI · Anthropic · Google GenAI · LangChain + LangGraph · LiteLLM · Strands · Agno · CrewAI · PydanticAI · AutoGen · LlamaIndex · FastAPI · MCP

Straight talk: a runtime left in record mode measures and changes nothing, whichever mode the client asks for, so set both. Decision reports say what was replaced and why, and carry no token counters; provider usage is the only savings number that counts. Runtime unreachable means your original request goes through untouched, unless you opt into strict mode.

Docs: middleware overview · Vercel AI SDK guide · Python guide · every framework and version · deploy beside your app · package READMEs for TypeScript and Python.

Rather not touch code? Point any SDK at the proxy with a baseURL swap instead: integrations/recipes/.


🧭 When to use · when to skip

Good fit if you read your agent's answers more than you paste them somewhere, run long sessions full of logs and test output, or pay per token and want the reading side shrunk without changing your code.

Skip it if you are billed per request rather than per token (GitHub Copilot premium requests, for one: a shorter answer is the same request), or your workload is pure code generation with almost no prose to cut. The ruleset rides along as input tokens on every call (about 1,000 estimated for the full skill), and on terse one-liner Q&A that can cost more than it saves.

Measure it yourself. Run the same task with and without caveman and compare the provider's billing page. That A/B outranks every number on this page. If caveman loses on your workload, turn it off. Full list of where it loses: docs/HONEST-NUMBERS.md.


🏔 The whole cave

One idea everywhere: agent do more with less.

Repo What it shrinks Status
caveman (you here) What the agent says (skill), reads (proxy), and what your own app sends (middleware) live
caveman-browse What the agent sees in the browser live
caveman-agent-sdk What your production agent loads, calls, and spends own repo · in dev
cavegemma The compression baked into weights (Gemma fine-tune) labs
caveman-code The whole agent, end to end frozen
cavemem What the agent remembers, across sessions frozen
cavekit The build loop, spec-driven frozen

Frozen ones still install and work. Their best ideas moved in here.

Caveman make token small. Caveman Cloud make it provable. Local numbers are inferred, pinned benchmarks benchmark_counterfactual, neither is an invoice. Live traffic behind eval gates with signed receipts earns verified. That's Cloud. Waitlist at caveman.so


🔒 Privacy, and a small favor

Your agent still talks to the provider you chose. The skill runs entirely on your machine, and nothing here needs an account.

The caveman CLI does send usage stats by default, and here's the honest why: caveman is free, one person maintains it, and those stats are how I find out which commands people actually use and which optimizations run in real workflows. That's what keeps this thing free and pointed in the right direction. Fair trade, we think.

What it sends: which commands ran, when your agent starts a session (the CLI's agent hooks send that one), token counts through and cut, a random install ID, your OS and CLI version, whether you're signed in, how you installed it, your timezone and language, and the IP address the stats come from. IPs get wiped after 90 days, everything else after 13 months. What it never sends: your prompts, your code, or your file paths. It tells you all this the first time you run it.

Not into it? One command and it's off forever, no hard feelings:

caveman telemetry off      # or set DO_NOT_TRACK=1

Want what it already sent gone too? telemetry off prints your install ID one last time. Send it to us and we delete it all. How: SECURITY.md.

Exact network, telemetry, and storage boundaries: SECURITY.md.


📜 License

One license: Apache-2.0, whole repo, from Caveman 3.0.0 on. Skill, CLI, client SDKs, middleware, contracts, provider catalog, extension, and the full runtime: Engine, Proxy, Browse, MCP server, shrink, cavemem, shared Go platform. Read it, fork it, ship it, host it. Free like mammoth on open plain.

Releases before 3.0.0 keep the license they shipped with. Details in LICENSING.md.

engine/pixel embeds pxpipe (MIT) plus glyph atlases derived from Spleen 5×8 (BSD-2-Clause) and GNU Unifont (dual OFL-1.1 / GPLv2-with-font-exception); its NOTICE travels with that source.

"Caveman" and the rock logo are trademarks of Julius Brussee. "Powered by Caveman" is fine when true.

📚 Cite

If caveman shows up in your paper, the way it showed up in Adobe's:

@software{brussee2026caveman,
  author = {Brussee, Julius},
  title  = {Caveman: why use many token when few do trick},
  year   = {2026},
  url    = {https://github.com/JuliusBrussee/caveman}
}

⭐ Star this repo

Caveman save you token, save you money. Star cost zero. Fair trade. ⭐


Docs: Technical manual · Install matrix · Honest numbers · Wrap benchmark · License · Contributing · Maintainer guide · Issues
One license, Apache-2.0. Few token. No lie.
View on GitHub

Recent activity

commits and pull requests

Releases and announcements

42 total
  1. Caveman 3.0.0v3.0.0Sep 30, 2026

    # Caveman 3.0.0 ## Everything is Apache-2.0 now The whole repo, including engine, proxy, browse, MCP server, shrink, the cavemem Go core and shared platform, is licensed Apache-2.0 from 3.0.0 onward. There's no hosted-service restriction and no Change Date, and you don't need a commercial license. Fork it, embed it, run it for your company. [LICENSING.md](https://github.com/JuliusBrussee/caveman/blob/v3.0.0/LICENSING.md) has the details. Earlier releases keep the license they shipped with. Code that outside contributors wrote under MIT stays available under MIT (`LICENSE-MIT`). ## Caveman Learn runs itself - **Autopilot.** After a session ends, Learn re-checks your history in the background at low priority, at most every 6 hours. It never makes a session wait, and `caveman learn autopilot off` stops it. - **Memory doctor.** It finds `MEMORY.md` notes Claude Code never loads (Claude Code only reads the first 200 lines or 25KB), `@imports` pointing at missing files, the same rule loaded from two files, and stale paths in `CLAUDE.md`. - **Trends.** Week-over-week numbers from your own sessions. They use the median session, and you get "not enough data" when there are fewer than 5

  2. @caveman-ai/sdk 1.2.0sdk-ts-v1.2.0Sep 30, 20264 downloads

    - **Breaking (license):** relicensed from MIT to Apache-2.0, along with the rest of the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - `onDecision` and `onDiagnostic` accept async sinks: a rejected promise is swallowed like a throw, as `onReport` already was, instead of an unhandled rejection that terminates Node. - Callers that only waited on another call's shared capabilities fetch no longer record its failure in the breaker: one refused connect at cold start is one failure, not one per waiter. A call its host aborts records nothing (it used to record a success, so a cancelled half-open probe closed the breaker) and frees the probe slot. New `CircuitBreaker.release()`. - A runtime token is trimmed of surrounding whitespace; any other character outside printable ASCII is `invalid_configuration`, as in Python. - An `https://` or `socks` `HTTPS_PROXY`/`HTTP_PROXY` is `invalid_configuration` with the unsupported scheme named by `ready()` and `preflight()`, instead of `runtime_unavailable` on every call. `MiddlewareError` takes an optional detail. - `retrieve()` (and so the recovery binding) throws `MiddlewareError('invalid_request

  3. caveman-sdk 1.2.0sdk-python-v1.2.0Sep 30, 20263 downloads

    - **Breaking (license):** relicensed from MIT to Apache-2.0, along with the rest of the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - Callers that only waited on another call's shared capabilities fetch no longer record its failure in the breaker: one refused connect at cold start is one failure, not one per waiter. An interrupted half-open probe records nothing and frees the probe slot. New `CircuitBreaker.release()`. - `observe()`/`observe_background()` normalize the receipt scope, as TypeScript does, and drop a receipt with an invalid scope. Async `observe()` has its own receipt worker (one worker, sixteen queued), so receipts never take optimize's slots. - A runtime token is stripped of surrounding whitespace (a mounted secret file's trailing newline); a token with any other character outside printable ASCII is `invalid_configuration` instead of a `runtime_unavailable` on every call. - An `https://` or `socks` proxy URL is `invalid_configuration`; `ready()` and `preflight()` name the unsupported scheme. `MiddlewareError` takes an optional detail. - `endpoint` is the runtime origin (`scheme://host[:port]`), as in TypeScript,

  4. @caveman-ai/middleware 1.0.0middleware-ts-v1.0.0Sep 30, 20262 downloads

    - First stable release. Requires `@caveman-ai/sdk` 1.2.0 or later. - **Breaking (license):** relicensed from MIT to Apache-2.0, along with the rest of the repository in Caveman 3.0.0. Releases before this one keep the MIT license. - Release process: prereleases no longer take the `latest` dist-tag, each release gets a GitHub Release with these notes and a CycloneDX SBOM, and the published dependency graph is audited before publish. - Version gate: - A framework version outside the tested range, or a prerelease, now passes content through with one `unsupported_version` warning instead of silently doing nothing. `acceptFrameworkVersion` overrides it. - Bundled deploys run with a one-time `version_unverified` notice. - Nothing throws at wrap time; strict mode raises from `ready()`. - The framework version is read from the application's installed copy. Framework peers are declared optional and unranged (`*`), so Yarn PnP can resolve them without a plain `npm install` failing with ERESOLVE. - `require()` works, and TypeScript resolves with node10, node16, nodenext and bundler. `./compatibility` exposes a `tier`. Importing under `workerd`/`edge-light` throws

  5. caveman-middleware 1.0.0middleware-python-v1.0.0Sep 30, 20264 downloads

    - First stable release. Requires `caveman-sdk` 1.2.0 or later (`caveman-sdk>=1.2,<2`): 1.1.0 lacks APIs the adapters import. - `caveman_middleware.__version__` is read from the installed distribution. - A refused `caveman_retrieve` (an unknown or expired handle, or a 404/410/503) returns `{"error": code}` through each framework's tool-error result and warns once. The run carries on, and cancellation still propagates. - Recovery arguments that are not an object with a string `handle` (`None`, a list) return `{"error": "invalid_request"}` in the OpenAI, Anthropic and MCP adapters. - Wrapping a client, agent or model twice runs one Caveman layer instead of stacking middleware and turning compression off. - A wrapped `AnthropicBedrock` client keeps `aws_profile`, so calls are signed with the caller's AWS identity. - LiteLLM registers one process-wide callback however many instances are live, so the host's own callbacks are never dropped. - Importing an adapter whose framework is out of range names the installed version and the supported range. - Every reported reason is a spec §8 catalog code. Calls that were never LLM calls (embeddings, token counting, non-LLM A

Code frequency

additions and deletions
+267.3K-267.3KWeek of 2026-03-29: +351 linesWeek of 2026-03-29: -25 linesWeek of 2026-04-05: +12,833 linesWeek of 2026-04-05: -5,904 linesWeek of 2026-04-12: +1,382 linesWeek of 2026-04-12: -198 linesWeek of 2026-04-19: +0 linesWeek of 2026-04-19: -0 linesWeek of 2026-04-26: +5,018 linesWeek of 2026-04-26: -810 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +8,828 linesWeek of 2026-05-10: -7,464 linesWeek of 2026-05-17: +45 linesWeek of 2026-05-17: -20 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +1,395 linesWeek of 2026-05-31: -165 linesWeek of 2026-06-07: +246 linesWeek of 2026-06-07: -378 linesWeek of 2026-06-14: +619 linesWeek of 2026-06-14: -18 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +2,099 linesWeek of 2026-06-28: -443 linesWeek of 2026-07-05: +0 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +0 linesWeek of 2026-07-12: -0 linesWeek of 2026-07-19: +4,197 linesWeek of 2026-07-19: -2,590 linesWeek of 2026-07-26: +343 linesWeek of 2026-07-26: -9 linesWeek of 2026-08-02: +541 linesWeek of 2026-08-02: -156 linesWeek of 2026-08-09: +267,280 linesWeek of 2026-08-09: -3,959 linesWeek of 2026-08-16: +24,586 linesWeek of 2026-08-16: -2,712 linesWeek of 2026-08-23: +10,738 linesWeek of 2026-08-23: -12,074 linesWeek of 2026-08-30: +14,075 linesWeek of 2026-08-30: -12,955 linesWeek of 2026-09-06: +19,785 linesWeek of 2026-09-06: -2,850 linesWeek of 2026-09-13: +182,162 linesWeek of 2026-09-13: -156,841 linesWeek of 2026-09-20: +46,855 linesWeek of 2026-09-20: -57,767 linesWeek of 2026-09-27: +8,192 linesWeek of 2026-09-27: -1,042 linesMar 29, 2026Sep 27, 2026
+611.6K lines added, -268.4K removed over the last year.

Commits per week

last 52 weeks
1000Week of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 6 commitsWeek of 2026-04-05: 81 commitsWeek of 2026-04-12: 20 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 13 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 20 commitsWeek of 2026-05-17: 3 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 9 commitsWeek of 2026-06-07: 9 commitsWeek of 2026-06-14: 5 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 27 commitsWeek of 2026-07-05: 0 commitsWeek of 2026-07-12: 0 commitsWeek of 2026-07-19: 7 commitsWeek of 2026-07-26: 2 commitsWeek of 2026-08-02: 12 commitsWeek of 2026-08-09: 31 commitsWeek of 2026-08-16: 80 commitsWeek of 2026-08-23: 53 commitsWeek of 2026-08-30: 100 commitsWeek of 2026-09-06: 57 commitsWeek of 2026-09-13: 82 commitsWeek of 2026-09-20: 37 commitsWeek of 2026-09-27: 61 commitsOct 5, 2025Sep 27, 2026
715 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 6 commitsSun 1:00 — 14 commitsSun 2:00 — 7 commitsSun 3:00 — 7 commitsSun 4:00 — 0 commitsSun 5:00 — 12 commitsSun 6:00 — 9 commitsSun 7:00 — 4 commitsSun 8:00 — 2 commitsSun 9:00 — 3 commitsSun 10:00 — 22 commitsSun 11:00 — 2 commitsSun 12:00 — 5 commitsSun 13:00 — 16 commitsSun 14:00 — 5 commitsSun 15:00 — 4 commitsSun 16:00 — 2 commitsSun 17:00 — 21 commitsSun 18:00 — 16 commitsSun 19:00 — 4 commitsSun 20:00 — 3 commitsSun 21:00 — 9 commitsSun 22:00 — 3 commitsSun 23:00 — 1 commitsMon 0:00 — 1 commitsMon 1:00 — 2 commitsMon 2:00 — 6 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 1 commitsMon 7:00 — 14 commitsMon 8:00 — 1 commitsMon 9:00 — 14 commitsMon 10:00 — 5 commitsMon 11:00 — 0 commitsMon 12:00 — 2 commitsMon 13:00 — 0 commitsMon 14:00 — 1 commitsMon 15:00 — 17 commitsMon 16:00 — 7 commitsMon 17:00 — 8 commitsMon 18:00 — 0 commitsMon 19:00 — 1 commitsMon 20:00 — 10 commitsMon 21:00 — 11 commitsMon 22:00 — 0 commitsMon 23:00 — 1 commitsTue 0:00 — 1 commitsTue 1:00 — 3 commitsTue 2:00 — 7 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 1 commitsTue 8:00 — 0 commitsTue 9:00 — 2 commitsTue 10:00 — 19 commitsTue 11:00 — 24 commitsTue 12:00 — 2 commitsTue 13:00 — 3 commitsTue 14:00 — 10 commitsTue 15:00 — 3 commitsTue 16:00 — 1 commitsTue 17:00 — 5 commitsTue 18:00 — 5 commitsTue 19:00 — 1 commitsTue 20:00 — 1 commitsTue 21:00 — 6 commitsTue 22:00 — 1 commitsTue 23:00 — 10 commitsWed 0:00 — 6 commitsWed 1:00 — 3 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 1 commitsWed 6:00 — 0 commitsWed 7:00 — 9 commitsWed 8:00 — 2 commitsWed 9:00 — 1 commitsWed 10:00 — 9 commitsWed 11:00 — 9 commitsWed 12:00 — 4 commitsWed 13:00 — 4 commitsWed 14:00 — 10 commitsWed 15:00 — 14 commitsWed 16:00 — 7 commitsWed 17:00 — 0 commitsWed 18:00 — 14 commitsWed 19:00 — 2 commitsWed 20:00 — 1 commitsWed 21:00 — 1 commitsWed 22:00 — 3 commitsWed 23:00 — 8 commitsThu 0:00 — 4 commitsThu 1:00 — 12 commitsThu 2:00 — 4 commitsThu 3:00 — 0 commitsThu 4:00 — 11 commitsThu 5:00 — 0 commitsThu 6:00 — 1 commitsThu 7:00 — 7 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 0 commitsThu 12:00 — 2 commitsThu 13:00 — 9 commitsThu 14:00 — 8 commitsThu 15:00 — 12 commitsThu 16:00 — 3 commitsThu 17:00 — 11 commitsThu 18:00 — 2 commitsThu 19:00 — 5 commitsThu 20:00 — 13 commitsThu 21:00 — 8 commitsThu 22:00 — 0 commitsThu 23:00 — 2 commitsFri 0:00 — 9 commitsFri 1:00 — 11 commitsFri 2:00 — 6 commitsFri 3:00 — 0 commitsFri 4:00 — 1 commitsFri 5:00 — 4 commitsFri 6:00 — 0 commitsFri 7:00 — 5 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 0 commitsFri 11:00 — 1 commitsFri 12:00 — 3 commitsFri 13:00 — 1 commitsFri 14:00 — 6 commitsFri 15:00 — 1 commitsFri 16:00 — 0 commitsFri 17:00 — 0 commitsFri 18:00 — 0 commitsFri 19:00 — 1 commitsFri 20:00 — 3 commitsFri 21:00 — 0 commitsFri 22:00 — 1 commitsFri 23:00 — 0 commitsSat 0:00 — 2 commitsSat 1:00 — 1 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 1 commitsSat 8:00 — 0 commitsSat 9:00 — 1 commitsSat 10:00 — 2 commitsSat 11:00 — 6 commitsSat 12:00 — 9 commitsSat 13:00 — 2 commitsSat 14:00 — 0 commitsSat 15:00 — 3 commitsSat 16:00 — 10 commitsSat 17:00 — 3 commitsSat 18:00 — 6 commitsSat 19:00 — 3 commitsSat 20:00 — 1 commitsSat 21:00 — 1 commitsSat 22:00 — 8 commitsSat 23:00 — 2 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits613 (74%)
Community commits210 (26%)

823 commits in total over the last year.

DateListRankStars gained
Jul 15, 2026daily#14+7
Jul 14, 2026daily#16+5
Jul 13, 2026daily#8+4
Jul 12, 2026daily#12+3
Jul 5, 2026daily#7+31
Jul 4, 2026daily#1+42
Jul 3, 2026daily#2+32
Jul 2, 2026daily#4+43
May 11, 2026daily#24+92
May 7, 2026daily#17+57
May 6, 2026daily#25+57
May 5, 2026daily#25+53
May 1, 2026daily#15+69
Apr 30, 2026daily#15+53
Apr 29, 2026daily#12+90
  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript

  • firecrawl/firecrawl

    Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥

    188.6K stars · TypeScript