davebcn87/pi-autoresearchPublic

Autonomous experiment loop extension for pi

AI summary: An autonomous experimentation loop that continually modifies code, runs benchmarks, and persists improvements.

Stars
8.1K
+7 today
Forks
465
Watchers
27
Open issues
4
Open PRs
6
Contributors
~18
Commits
174
Branches
12

TypeScriptMITCreated Mar 11, 2026Last push 24d agoLatest release v1.8.1+31 stars this week+157 this month

Quick answers

What is pi-autoresearch?
An autonomous experimentation loop that continually modifies code, runs benchmarks, and persists improvements.
What does pi-autoresearch do?
pi-autoresearch orchestrates an autonomous optimization loop for the 'pi' coding agent by continuously editing code, committing changes, and running benchmarks. It relies on a provided bash script to measure metrics and records every iteration in an append-only JSONL log, preserving history across agent restarts or context window limits. The system calculates a robust confidence score using Median Absolute Deviation to distinguish genuine improvements from benchmark noise. Once an optimization session concludes, it groups successful iterations into logically independent, reviewable branches starting from the merge-base, ensuring they can be merged without conflict.
Who is pi-autoresearch for?
This tool is for software engineers, ML researchers, and developers looking to automate the tedious process of performance optimization and hyperparameter tuning. It requires the 'pi' coding agent environment and basic shell scripting knowledge to write measurement scripts.
How do I get started with pi-autoresearch?
pi install npm:pi-autoresearch
How popular is pi-autoresearch on GitHub?
davebcn87/pi-autoresearch has 8,148 stars and 465 forks on GitHub, and gained 31 stars in the last 7 days.
What license does pi-autoresearch use?
davebcn87/pi-autoresearch is released under the MIT license.

Star history

since Jul 29, 2026
02.5K5K7.5KJul 2026Aug 2026Sep 2026Oct 2026
8.1K stars as of Oct 4, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 6 commits2026-03-11: 54 commits2026-03-12: 4 commits2026-03-13: 5 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 8 commits2026-03-17: 3 commits2026-03-18: 13 commits2026-03-19: 0 commits2026-03-20: 2 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 5 commits2026-03-26: 6 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 1 commit2026-04-01: 0 commits2026-04-02: 1 commit2026-04-03: 0 commits2026-04-04: 1 commit2026-04-05: 0 commits2026-04-06: 2 commits2026-04-07: 0 commits2026-04-08: 1 commit2026-04-09: 1 commit2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 1 commit2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 1 commit2026-04-21: 1 commit2026-04-22: 2 commits2026-04-23: 1 commit2026-04-24: 3 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 7 commits2026-04-29: 2 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 2 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 2 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 4 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 2 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 3 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 1 commit2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 0 commits2026-08-13: 0 commits2026-08-14: 0 commits2026-08-15: 0 commits2026-08-16: 0 commits2026-08-17: 0 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 2 commits2026-09-01: 2 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 4 commits2026-09-09: 0 commits2026-09-10: 1 commit2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits
154 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

What pi-autoresearch does

pi-autoresearch orchestrates an autonomous optimization loop for the 'pi' coding agent by continuously editing code, committing changes, and running benchmarks. It relies on a provided bash script to measure metrics and records every iteration in an append-only JSONL log, preserving history across agent restarts or context window limits. The system calculates a robust confidence score using Median Absolute Deviation to distinguish genuine improvements from benchmark noise. Once an optimization session concludes, it groups successful iterations into logically independent, reviewable branches starting from the merge-base, ensuring they can be merged without conflict.

This tool is for software engineers, ML researchers, and developers looking to automate the tedious process of performance optimization and hyperparameter tuning. It requires the 'pi' coding agent environment and basic shell scripting knowledge to write measurement scripts.

  • Autonomous Optimization Loop: Continually modifies code and runs a benchmark command to find improvements until interrupted.
  • Persistent State Logging: Writes all runs to an append-only JSONL file and maintains a living prompt to survive context resets.
  • Statistical Confidence Scoring: Uses Median Absolute Deviation to evaluate whether metric improvements are real or just benchmark jitter.
  • Branch Finalization: Automatically regroups successful experiments into isolated, reviewable branches based on the files they touch.
  • Backpressure Checks: Supports optional scripts to run test suites or linters after passing benchmarks to prevent breaking correctness.
  • Real-time Dashboard: Provides a terminal widget, fullscreen overlay, and browser export to monitor optimization progress live.

Where teams use it

Performance Tuning

Developers use it to blindly search for optimizations to reduce bundle sizes or speed up build times.

Machine Learning Optimization

AI researchers let the tool continually tweak training scripts to lower validation loss over long unattended sessions.

Test Suite Acceleration

Engineers employ it to experiment with test runner configurations and parallelization to decrease CI runtime.

Lighthouse Score Improvement

Frontend developers run it overnight to discover code changes that incrementally boost web performance metrics.

Getting started: pi install npm:pi-autoresearch

README

main branch
result

pi-autoresearch

Autonomous experiment loop for pi

Website · Documentation · Install · Usage · How it works

Try an idea, measure it, keep what works, discard what doesn't, repeat forever.

An extension for pi — an AI coding agent that runs in your terminal. pi-autoresearch gives pi the tools and workflow to run autonomous optimization loops: try an idea, benchmark it, keep improvements, revert regressions, repeat.

Inspired by karpathy/autoresearch. Works for any optimization target: test speed, bundle size, LLM training, build times, Lighthouse scores.


pi-autoresearch

Quick start

pi install npm:pi-autoresearch
pi

Or load the package for one session only with pi -e npm:pi-autoresearch.

Then start the loop inside pi:

/autoresearch optimize unit test runtime, monitor correctness

What's included

Extension Tools + live widget + /autoresearch dashboard
Skill Gathers what to optimize, writes session files, starts the loop

Extension tools

Tool Description
init_experiment One-time session config — name, metric, unit, direction
run_experiment Runs any command, times wall-clock duration, captures output
log_experiment Records result, auto-commits, updates widget and dashboard

/autoresearch command

Subcommand Description
/autoresearch Show help without activating autoresearch mode.
/autoresearch <text> Enter autoresearch mode. If .auto/prompt.md exists, resumes the loop with <text> as context. Otherwise, sets up a new session.
/autoresearch off Leave autoresearch mode. Stops auto-resume and clears runtime state but keeps .auto/log.jsonl intact.
/autoresearch clear Delete .auto/log.jsonl, reset all state, and turn autoresearch mode off. Use this for a clean start.
/autoresearch export Open a live dashboard in your browser. Auto-updates as experiments run.
/autoresearch dashboard Open the fullscreen scrollable dashboard overlay in the terminal. Navigate with ↑/↓/j/k, PageUp/PageDown/u/d, g/G for top/bottom, Escape or q to close.

Examples:

/autoresearch optimize unit test runtime, monitor correctness
/autoresearch model training, run 5 minutes of train.py and note the loss ratio as optimization target
/autoresearch export
/autoresearch dashboard
/autoresearch off
/autoresearch clear

Keyboard shortcuts (opt-in)

No shortcuts are bound by default — pi's built-in keymap grows with every release, so any default chord eventually collides with it (this happened with ctrl+shift+f vs. pi's transcript search). Every action is available as a /autoresearch subcommand instead.

To opt in, add chords to <agent-dir>/extensions/pi-autoresearch.json (<agent-dir> is usually ~/.pi/agent, or PI_CODING_AGENT_DIR when set). Omitted or null keys stay unbound:

{
  "shortcuts": {
    "fullscreenDashboard": "ctrl+shift+y",
    "export": "alt+shift+e",
    "off": null
  }
}

Each key binds its /autoresearch subcommand: fullscreenDashboard → dashboard, export → export, off → off.

Pick a chord that's free in pi's built-in keybindings and your <agent-dir>/keybindings.json — extension shortcuts win conflicts, so a clashing chord hijacks the built-in action.

Picking a free chord (guidance for agents)

Don't guess a chord — check it against the installed pi's built-in keymap first:

PI_ROOT=$(ls -d ~/.pi/pkg/pi-* 2>/dev/null | sort -V | tail -1)
[ -n "$PI_ROOT" ] || PI_ROOT="$(npm root -g)/@earendil-works/pi-coding-agent"
node --input-type=module -e '
const { KEYBINDINGS } = await import(process.argv[1] + "/dist/core/keybindings.js");
const taken = new Set(Object.values(KEYBINDINGS).flatMap(({ defaultKeys }) => defaultKeys).map((key) => key.toLowerCase()));
console.log(taken.has(process.argv[2].toLowerCase()) ? "TAKEN" : "FREE");
' "$PI_ROOT" ctrl+shift+y

Chords in <agent-dir>/keybindings.json values are also taken. Write a verified-free chord into <agent-dir>/extensions/pi-autoresearch.json, then /reload pi and confirm no Extension shortcut conflict warning appears.

ctrl+shift+y, ctrl+shift+u, alt+shift+f, and ctrl+alt+d are free as of pi 0.84.x. Always re-verify — pi's keymap grows over time, which is why this extension binds nothing by default.

UI

  • Dashboard widget — always visible above the editor: a full results table with columns for commit, metric, status, and description.
  • Confidence score — after 3+ runs, shows how the best improvement compares to the session noise floor. ≥2.0× (green) = likely real, 1.0–2.0× (yellow) = above noise but marginal, <1.0× (red) = within noise.
  • Fullscreen overlay — /autoresearch dashboard opens a scrollable full-terminal dashboard. Shows a live spinner with elapsed time for running experiments.

Skills

autoresearch-create asks a few questions (or infers from context) about your goal, command, metric, and files in scope — then writes two files and starts the loop immediately:

autoresearch-finalize turns a noisy autoresearch branch into clean, independent branches — one per logical change, each starting from the merge-base. Groups must not share files, so each branch can be reviewed and merged independently.

autoresearch-hooks (optional) helps author .auto/hooks/before.sh and .auto/hooks/after.sh for a session. It ships with ten reference scripts in skills/autoresearch-hooks/examples/ (external search, learnings journal, native notifications, anti-thrash, idea rotation, and more) — the skill handles the contract, you pick the inspiration. The core autoresearch loop has no hook awareness.

All session files live in a single .auto/ subfolder at the working-directory root — one folder to preserve across reverts, gitignore, and clean up. (Legacy flat autoresearch.* files are still read for in-flight sessions.)

File Purpose
.auto/prompt.md Session document — objective, metrics, files in scope, what's been tried. A fresh agent can resume from this alone.
.auto/measure.sh Benchmark script — pre-checks, runs the workload, outputs METRIC name=number lines.
.auto/log.jsonl Append-only log of every run (written by the tools).
.auto/checks.sh (optional) Backpressure checks — tests, types, lint. Runs after each passing benchmark. Failures block keep.
.auto/hooks/ (optional) Executable scripts (before.sh, after.sh) that fire around iterations. Stdout is delivered to the agent as a steer message.

Install

pi install npm:pi-autoresearch
Manual install
cp -r extensions/pi-autoresearch ~/.pi/agent/extensions/
cp -r skills/autoresearch-create ~/.pi/agent/skills/

Then /reload in pi.


Usage

1. Start autoresearch

/autoresearch optimize unit test runtime, monitor correctness

This activates autoresearch mode and makes the experiment tools available. If .auto/prompt.md does not exist, it also loads the autoresearch-create skill. Calling /skill:autoresearch-create directly does not activate the mode, so the tools remain unavailable in a fresh session.

The agent asks about your goal, command, metric, and files in scope — or infers them from context. It then creates a branch, writes .auto/prompt.md and .auto/measure.sh, runs the baseline, and starts looping immediately.

2. The loop

The agent runs autonomously: edit → commit → run_experiment → log_experiment → keep or revert → repeat. It never stops unless interrupted.

Every result is appended to .auto/log.jsonl in your project — one line per run. This means:

  • Survives restarts — the agent can resume a session by reading the file
  • Survives context resets — .auto/prompt.md captures what's been tried so a fresh agent has full context
  • Human readable — open it anytime to see the full history
  • Branch-aware — each branch has its own session

3. Finalize into reviewable branches

/skill:autoresearch-finalize

The agent reads .auto/log.jsonl, groups kept experiments into logical changesets, proposes the grouping for your approval, then creates independent branches from the merge-base. Each commit includes metric improvements in the message. Groups must not share files, so branches can be reviewed and merged independently.

4. Monitor progress

  • Widget — full results table, always visible above the editor
  • /autoresearch dashboard — fullscreen scrollable dashboard overlay (optional chord: shortcuts.fullscreenDashboard)
  • /autoresearch export — open a live browser dashboard with chart and share card
  • Escape — interrupt anytime and ask for a summary

Example domains

Domain Metric Command
Test speed seconds ↓ pnpm test
Bundle size KB ↓ pnpm build && du -sb dist
LLM training val_bpb ↓ uv run train.py
Build speed seconds ↓ pnpm build
Lighthouse perf score ↑ lighthouse http://localhost:3000 --output=json

How it works

The extension is domain-agnostic infrastructure. The skill encodes domain knowledge. This separation means one extension serves unlimited domains.

┌──────────────────────┐     ┌──────────────────────────┐
│  Extension (global)  │     │  Skill (per-domain)       │
│                      │     │                           │
│  run_experiment      │◄────│  command: pnpm test       │
│  log_experiment      │     │  metric: seconds (lower)  │
│  widget + dashboard  │     │  scope: vitest configs    │
│                      │     │  ideas: pool, parallel…   │
└──────────────────────┘     └──────────────────────────┘

Two files keep the session alive across restarts and context resets:

.auto/log.jsonl   — append-only log of every run (metric, status, commit, description)
.auto/prompt.md   — living document: objective, what's been tried, dead ends, key wins

A fresh agent with no memory can read these two files and continue exactly where the previous session left off.


Configuration (optional)

Create .auto/config.json in your pi session directory to customize behavior:

{
  "workingDir": "/path/to/project",
  "maxIterations": 50
}
Field Type Description
workingDir string Override the directory for all autoresearch operations — file I/O, command execution, and git. Supports absolute or relative paths (resolved against the pi session cwd). The config file itself always stays under the session cwd. Fails if the directory doesn't exist.
maxIterations number Maximum experiments before auto-stopping. The agent is told to stop and won't run more experiments until a new segment is initialized.

Long-running loops and context

The loop is designed to run unattended across context limits. When pi's auto-compaction summarizes the older portion of the conversation, autoresearch detects the resulting idle and re-prompts the agent to re-read .auto/prompt.md, the tail of .auto/log.jsonl, .auto/ideas.md, and git log before continuing. All progress is persisted in those files, so the post-summary turn rehydrates from the source of truth instead of relying on whatever survived compaction. No tuning required — if pi's auto-compaction is enabled (the default), this just works.


Confidence scoring

After 3+ experiments in a session, pi-autoresearch computes a confidence score — how the best improvement compares to the session's noise floor. This helps distinguish real gains from benchmark jitter, especially on noisy signals like ML training, Lighthouse scores, or flaky benchmarks.

How it works:

  • Uses Median Absolute Deviation (MAD) of all metric values in the current segment as a robust noise estimator.
  • Confidence = |best_improvement| / MAD. A score of 2.0× means the best improvement is twice the noise floor.
  • Shown in the widget, fullscreen dashboard, and log_experiment output.
  • Persisted to .auto/log.jsonl on each result for post-hoc analysis.
  • Advisory only — never auto-discards. The agent is guided to re-run experiments when confidence is low, but the final keep/discard decision stays with the agent.
Confidence Color Meaning
≥ 2.0× 🟢 green Improvement is likely real
1.0–2.0× 🟡 yellow Above noise but marginal
< 1.0× 🔴 red Within noise — consider re-running to confirm

Backpressure checks (optional)

Create .auto/checks.sh to run correctness checks (tests, types, lint) after every passing benchmark. This ensures optimizations don't break things.

#!/bin/bash
set -euo pipefail
pnpm test --run
pnpm typecheck

How it works:

  • If the file doesn't exist, everything behaves exactly as before — no changes to the loop.
  • If it exists, it runs automatically after every benchmark that exits 0.
  • Checks execution time does not affect the primary metric.
  • If checks fail, the experiment is logged as checks_failed (same behavior as a crash — no commit, revert changes).
  • The checks_failed status is shown separately in the dashboard so you can distinguish correctness failures from benchmark crashes.
  • Checks have a separate timeout (default 300s, configurable via checks_timeout_seconds in run_experiment).

Hooks (optional)

Drop executable scripts in .auto/hooks/ to run code at iteration boundaries. Hooks are transparent to the agent — the agent calls tools and sees results; hooks run alongside without any agent-facing surface.

  • .auto/hooks/before.sh — fires before every iteration (at /autoresearch activation and at the end of every log_experiment, after after.sh). Use for prospective work: fetch research, prime context for the next attempt.
  • .auto/hooks/after.sh — fires at the end of every log_experiment. Use for retrospective work: annotate learnings, send notifications.

Contract:

  • Must be executable (chmod +x). Preserved on revert — the entire .auto/ folder survives (as do legacy autoresearch.* artefacts).
  • Stdin — a JSON object on a single line. Shape depends on the stage (see below). Extract fields with jq.
  • Stdout is delivered to the agent as a steer message (capped at 8 KB). Empty stdout = silent.
  • Non-zero exit or >30s timeout surfaces an error steer to the agent.
  • Each fire appends a {"type":"hook",…} entry to .auto/log.jsonl for observability.

before.sh stdin (on fresh activation last_run is null):

{
  "event": "before",
  "cwd": "/path/to/workdir",
  "next_run": 6,
  "last_run": {
    "run": 5, "status": "discard", "metric": 42.1,
    "description": "…",
    "asi": { "hypothesis": "…", "next_focus": "…" }
  },
  "session": {
    "metric_name": "total_ms", "metric_unit": "ms", "direction": "lower",
    "baseline_metric": 40.7, "best_metric": 33.5,
    "run_count": 5, "goal": "optimize sort speed"
  }
}

after.sh stdin:

{
  "event": "after",
  "cwd": "/path/to/workdir",
  "run_entry": {
    "run": 6, "status": "discard", "metric": 38.9,
    "description": "…",
    "asi": { "hypothesis": "…", "learned": "…" }
  },
  "session": { "metric_name": "total_ms", "direction": "lower", "baseline_metric": 40.7, "best_metric": 33.5, "run_count": 6, "goal": "…" }
}

Agent signal. The agent writes description and asi.* fields in its log_experiment calls for its own future-self reasoning. The hook opportunistically mines whichever fields the agent naturally uses — asi.hypothesis, asi.next_focus, description, etc. There is no dedicated "hook input" field; the agent is unaware the hook exists.

Examples. Reference scripts for both stages live at skills/autoresearch-hooks/examples/ — external search, qmd document search, persistent learnings, native notifications, git tagging, anti-thrash, idea rotator, hypothesis reflection, context rotation. Copy one to your session's .auto/hooks/ directory, adapt, chmod +x.


Prerequisites

  1. Install pi — follow the instructions at pi.dev
  2. An API key for your preferred LLM provider (configured in pi)

Security and trust

Pi packages run with your full user permissions. Review this repository before installing it, and run autoresearch in a dedicated branch or worktree with a clean working tree.

Autoresearch intentionally edits files, creates and reverts commits, and executes the commands in .auto/measure.sh, .auto/checks.sh, and .auto/hooks/. Treat those files as executable code: review them before each session, keep credentials and sensitive files out of scope, and use a sandbox or restricted environment for untrusted projects.

For reproducible installs, pin a version you have reviewed, for example pi install npm:[email protected]. Published npm releases include provenance attestations.

Controlling costs

Autoresearch loops run autonomously and can burn through tokens. Two ways to cap spend:

  • API key limits — most providers let you set per-key or monthly budgets. Check your provider's dashboard.
  • maxIterations — cap experiments per session in .auto/config.json:
    {
      "maxIterations": 30
    }

License

MIT

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

14 total
  1. v1.8.1v1.8.1Sep 8, 2026

    ### Fixed - `/autoresearch <goal>` without a `.auto/prompt.md` sent the literal text `/skill:autoresearch-create …` to the model instead of the skill's contents, because `pi.sendUserMessage()` does not expand skill commands by default. Models would reply with things like `Unknown command: /skill:autoresearch-create` (#93). The kickoff is now sent with `expandPromptTemplates: true` (pi ≥ 0.84.2).

  2. v1.8.0v1.8.0Sep 8, 2026

    ### Added - After every logged experiment, `log_experiment` now asks the agent to check whether the latest result invalidates a previous discard's rollback reason before choosing the next experiment. Ideas discarded because "X was the bottleneck" get a second look once X stops being the bottleneck. - Intentional retries can be annotated with `asi.revisits_run: <run number>`; the transcript then shows a `↻ Revisiting #N` line under the logged result so a retry is distinguishable from the agent forgetting a failure. - The `.auto/prompt.md` template's "What's Been Tried" section now asks for the conditions that would justify revisiting a discarded idea, so that knowledge survives compaction.

  3. v1.7.0v1.7.0Aug 31, 2026

    ### Changed - **Breaking:** no keyboard shortcuts are bound by default anymore. The fullscreen dashboard default chord `ctrl+shift+f` collided with pi 0.84.2's new built-in transcript search (#86) — and any hardcoded default will eventually collide with a future pi built-in. Shortcuts are now strictly opt-in via `<agent-dir>/extensions/pi-autoresearch.json`. To restore the old behavior: `{ "shortcuts": { "fullscreenDashboard": "ctrl+shift+f" } }`. ### Added - `/autoresearch dashboard` subcommand opens the fullscreen dashboard overlay — the keyboard-free way to reach it. - The `export` (`/autoresearch export`) and `off` (`/autoresearch off`) actions can now be bound to opt-in shortcuts alongside `fullscreenDashboard`. - README guidance (aimed at agents) for verifying a chord against the installed pi's built-in keymap before writing it to the shortcut config.

  4. v1.6.2v1.6.2Jul 9, 2026

    ### Changed - Raised the autoresearch auto-resume ceiling from 20 to 200 turns, while adding a stuck-loop override that stops auto-resume after more than 20 consecutive `discard` or `crash` results in the current segment.

  5. v1.6.1v1.6.1Jul 2, 2026

    ### Fixed - Redirected `workingDir` logs no longer auto-activate autoresearch in unrelated pi sessions. - `/autoresearch off` now persists across `/tree`, compaction, and reloads: a manual off is recorded as a session activation decision and is no longer overridden just because `log.jsonl` still exists.

Code frequency

additions and deletions
+8.5K-8.5KWeek of 2026-03-08: +8,527 linesWeek of 2026-03-08: -2,231 linesWeek of 2026-03-15: +2,513 linesWeek of 2026-03-15: -205 linesWeek of 2026-03-22: +1,562 linesWeek of 2026-03-22: -1,357 linesWeek of 2026-03-29: +288 linesWeek of 2026-03-29: -213 linesWeek of 2026-04-05: +4,238 linesWeek of 2026-04-05: -7 linesWeek of 2026-04-12: +141 linesWeek of 2026-04-12: -39 linesWeek of 2026-04-19: +1,313 linesWeek of 2026-04-19: -187 linesWeek of 2026-04-26: +1,757 linesWeek of 2026-04-26: -1,585 linesWeek of 2026-05-03: +413 linesWeek of 2026-05-03: -57 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +0 linesWeek of 2026-05-17: -0 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +33 linesWeek of 2026-05-31: -10 linesWeek of 2026-06-07: +1,545 linesWeek of 2026-06-07: -4,480 linesWeek of 2026-06-14: +0 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +506 linesWeek of 2026-06-28: -6 linesWeek of 2026-07-05: +143 linesWeek of 2026-07-05: -18 linesWeek of 2026-07-12: +14 linesWeek of 2026-07-12: -11 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesWeek of 2026-08-09: +0 linesWeek of 2026-08-09: -0 linesWeek of 2026-08-16: +0 linesWeek of 2026-08-16: -0 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +5,203 linesWeek of 2026-08-30: -221 linesWeek of 2026-09-06: +550 linesWeek of 2026-09-06: -29 linesWeek of 2026-09-13: +0 linesWeek of 2026-09-13: -0 linesMar 8, 2026Sep 13, 2026
+28.7K lines added, -10.7K removed over the last year.

Commits per week

last 52 weeks
690Week of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 69 commitsWeek of 2026-03-15: 26 commitsWeek of 2026-03-22: 11 commitsWeek of 2026-03-29: 3 commitsWeek of 2026-04-05: 4 commitsWeek of 2026-04-12: 1 commitsWeek of 2026-04-19: 8 commitsWeek of 2026-04-26: 9 commitsWeek of 2026-05-03: 2 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 0 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 2 commitsWeek of 2026-06-07: 4 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 2 commitsWeek of 2026-07-05: 3 commitsWeek of 2026-07-12: 1 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsWeek of 2026-08-09: 0 commitsWeek of 2026-08-16: 0 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 4 commitsWeek of 2026-09-06: 5 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsSep 27, 2025Sep 20, 2026
154 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 0 commitsSun 14:00 — 0 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 1 commitsMon 7:00 — 0 commitsMon 8:00 — 2 commitsMon 9:00 — 0 commitsMon 10:00 — 2 commitsMon 11:00 — 2 commitsMon 12:00 — 1 commitsMon 13:00 — 1 commitsMon 14:00 — 2 commitsMon 15:00 — 3 commitsMon 16:00 — 0 commitsMon 17:00 — 3 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 0 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 1 commitsTue 10:00 — 1 commitsTue 11:00 — 7 commitsTue 12:00 — 2 commitsTue 13:00 — 0 commitsTue 14:00 — 2 commitsTue 15:00 — 4 commitsTue 16:00 — 2 commitsTue 17:00 — 4 commitsTue 18:00 — 0 commitsTue 19:00 — 1 commitsTue 20:00 — 0 commitsTue 21:00 — 0 commitsTue 22:00 — 0 commitsTue 23:00 — 0 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 17 commitsWed 8:00 — 11 commitsWed 9:00 — 7 commitsWed 10:00 — 10 commitsWed 11:00 — 5 commitsWed 12:00 — 12 commitsWed 13:00 — 5 commitsWed 14:00 — 0 commitsWed 15:00 — 2 commitsWed 16:00 — 3 commitsWed 17:00 — 0 commitsWed 18:00 — 1 commitsWed 19:00 — 0 commitsWed 20:00 — 0 commitsWed 21:00 — 6 commitsWed 22:00 — 1 commitsWed 23:00 — 0 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 3 commitsThu 9:00 — 3 commitsThu 10:00 — 0 commitsThu 11:00 — 5 commitsThu 12:00 — 4 commitsThu 13:00 — 0 commitsThu 14:00 — 0 commitsThu 15:00 — 0 commitsThu 16:00 — 3 commitsThu 17:00 — 0 commitsThu 18:00 — 0 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 1 commitsThu 23:00 — 1 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 0 commitsFri 10:00 — 2 commitsFri 11:00 — 2 commitsFri 12:00 — 2 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 2 commitsFri 17:00 — 2 commitsFri 18:00 — 0 commitsFri 19:00 — 0 commitsFri 20:00 — 1 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 1 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits99 (57%)
Community commits75 (43%)

174 commits in total over the last year.

DateListRankStars gained
Apr 16, 2026daily#15+122
Mar 13, 2026daily#21+155
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    456.7K stars · TypeScript

  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • anomalyco/opencode

    The open source coding agent.

    211.7K stars · TypeScript

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript

  • microsoft/vscode

    Visual Studio Code

    193.5K stars · TypeScript

  • firecrawl/firecrawl

    Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥

    188.6K stars · TypeScript