aiming-lab/AutoResearchClawPublic

Fully autonomous & self-evolving research from idea to paper. Chat an Idea. Get a Paper. 🦞

AI summary: An autonomous research framework utilizing human-AI collaboration and self-reinforcing mechanisms for scientific inquiry.

Stars
14K
+9 today
Forks
1.6K
Watchers
56
Open issues
1
Open PRs
8
Contributors
~54
Commits
281
Branches
1

PythonMITCreated Mar 15, 2026Last push 25d agoLatest release v0.5.0+43 stars this week+60 this month

Star history

since Mar 15, 2026
05K10KMar 2026May 2026Jun 2026Aug 2026
14K stars as of Aug 7, 2026, tracked back to Mar 15, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 2 commits2026-03-15: 35 commits2026-03-16: 32 commits2026-03-17: 35 commits2026-03-18: 11 commits2026-03-19: 16 commits2026-03-20: 8 commits2026-03-21: 2 commits2026-03-22: 2 commits2026-03-23: 0 commits2026-03-24: 2 commits2026-03-25: 2 commits2026-03-26: 3 commits2026-03-27: 3 commits2026-03-28: 0 commits2026-03-29: 3 commits2026-03-30: 4 commits2026-03-31: 0 commits2026-04-01: 9 commits2026-04-02: 1 commit2026-04-03: 2 commits2026-04-04: 1 commit2026-04-05: 0 commits2026-04-06: 2 commits2026-04-07: 1 commit2026-04-08: 7 commits2026-04-09: 0 commits2026-04-10: 4 commits2026-04-11: 1 commit2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 4 commits2026-04-21: 1 commit2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 1 commit2026-05-19: 1 commit2026-05-20: 0 commits2026-05-21: 1 commit2026-05-22: 2 commits2026-05-23: 1 commit2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 5 commits2026-05-28: 4 commits2026-05-29: 1 commit2026-05-30: 0 commits2026-05-31: 1 commit2026-06-01: 3 commits2026-06-02: 0 commits2026-06-03: 1 commit2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 1 commit2026-06-16: 1 commit2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 1 commit2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 2 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
219 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    13,973 stars

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    3 trending appearances

What AutoResearchClaw does

AutoResearchClaw provides an end-to-end framework for automating complex scientific research tasks. By orchestrating LLM agents with specialized claw tools, it handles everything from literature reviews to experimental design and data analysis. The system differentiates itself through a self-reinforcing loop, allowing it to evaluate its own outputs and iteratively improve its methodology. It also includes explicit human-in-the-loop checkpoints, ensuring that researchers can steer the autonomous process during critical decision-making phases.

This framework is designed for academic researchers, data scientists, and R&D teams who want to accelerate the scientific discovery process. It requires familiarity with Python and API integration for the underlying LLMs.

  • Self-Reinforcing Loop: Automatically evaluates and refines its own research outputs to improve quality over time.
  • Human-AI Collaboration: Implements checkpoints that allow human researchers to guide the AI at critical junctures.
  • Automated Literature Review: Synthesizes large volumes of academic papers to identify research gaps and trends.
  • Experimental Design Generation: Proposes methodologies and test structures based on synthesized hypotheses.
  • Modular Claw Tools: Utilizes specialized toolsets designed for distinct phases of the scientific method.
  • Autonomous Execution: Orchestrates multi-step research processes without constant manual intervention.

Where teams use it

Accelerated literature reviews

Academic researchers can deploy the system to autonomously digest thousands of papers and generate a comprehensive literature synthesis.

Hypothesis generation

Data scientists can use the tool to analyze existing datasets and formulate novel experimental hypotheses.

Methodology planning

Lab scientists can leverage the AI to draft detailed experimental procedures with human oversight.

Continuous research tracking

Research institutions can run the framework in the background to continuously track and summarize developments in a niche field.

Getting started: Please consult the repository for specific installation commands, as it typically requires cloning and setting up Python environments.

README

main branch

AutoResearchClaw Logo

Chat an Idea. Get a Paper. Autonomous, Collaborative & Self-Evolving.

Just chat with OpenClaw: "Research X" → done.

📄 Our paper is on arXiv — come read it! AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration

AutoResearchClaw Framework

arXiv ARC-Bench on Hugging Face MIT License Python 3.11+ 2699 Tests Passed GitHub OpenClaw Compatible Discord

🇨🇳 中文 · 🇯🇵 日本語 · 🇰🇷 한국어 · 🇫🇷 Français · 🇩🇪 Deutsch · 🇪🇸 Español · 🇧🇷 Português · 🇷🇺 Русский · 🇸🇦 العربية

🏆 Paper Showcase · 🧑‍✈️ Co-Pilot Guide · 📖 Integration Guide · 💬 Discord Community


Sample Paper 🏆 Generated Paper Showcase

8 papers across 8 domains — math, statistics, biology, computing, NLP, RL, vision, robustness — generated fully autonomously or with Human-in-the-Loop co-pilot guidance.

View Showcase

🧪 We're looking for testers! Try the pipeline with your own research idea — from any field — and tell us what you think. Your feedback directly shapes the next version. → Testing Guide | → 中文测试指南 | → 日本語テストガイド


🔥 News

  • [05/19/2026] v0.5.0Multi-Domain Experiment Agents + ARC-Bench — Two headline updates. (1) Domain-specialist execution agents: the experiment stage (Stages 10–13) now routes beyond the default ML sandbox to specialist agents per field — high-energy physics (ColliderAgent: Lagrangian → FeynRules → MadGraph5 → Delphes via the Magnus cloud), biology (COBRApy genome-scale metabolic modelling), and statistics (simulation-study agent), with a generic Docker executor covering chemistry/materials. The pipeline auto-selects the right executor from the research domain. (2) ARC-Bench: a 55-topic open-ended autonomous-research benchmark spanning ML (25), HEP (10), quantum (10), biology (7), and statistics (3) — each topic ships a manifest (research question + conditions + metrics + datasets) and a rubric for graded scoring, all under experiments/arc_bench/, and also released on 🤗 Hugging Face. → Domain Integration Guide
  • [04/01/2026] v0.4.0Human-in-the-Loop Co-Pilot System — AutoResearchClaw is no longer purely autonomous. New HITL system adds 6 intervention modes (full-auto, gate-only, checkpoint, step-by-step, co-pilot, custom), per-stage policies, and deep human-AI collaboration. Includes: Idea Workshop for hypothesis co-creation, Baseline Navigator for experiment design review, Paper Co-Writer for collaborative drafting, SmartPause (confidence-driven dynamic intervention), ALHF intervention learning, anti-hallucination claim verification, cost budget guardrails, pipeline branching for parallel hypothesis exploration, and CLI commands (attach/status/approve/reject/guide). → Full HITL Guide
  • [03/30/2026] Flexible Skill Loading — AutoResearchClaw now supports loading open-source and custom skills from any discipline to further enhance your research experience. 20 pre-loaded skills are included as ready-to-use references, covering scientific writing, experiment design, chemistry, biology, and more — including an A-Evolve agentic evolution skill contributed by the community. Load your own via researchclaw skills install or drop a SKILL.md into .claude/skills/. See Skills Library.
  • [03/22/2026] v0.3.2Cross-Platform Support + Major Stability — AutoResearchClaw now runs on any ACP-compatible agent backend (Claude Code, Codex CLI, Copilot CLI, Gemini CLI, Kimi CLI) and supports messaging platforms (Discord, Telegram, Lark, WeChat) via OpenClaw bridge. New CLI-agent code generation backend delegates Stages 10 & 13 to external CLI agents with budget control and timeout management. Also includes anti-fabrication system (VerifiedRegistry + experiment diagnosis & repair loop), 100+ bug fixes, modular executor refactoring, --resume auto-detection, LLM retry hardening, and community-reported fixes.
Earlier releases
  • [03/18/2026] v0.3.1OpenCode Beast Mode + Community Contributions — New "Beast Mode" routes complex code generation to OpenCode with automatic complexity scoring and graceful fallback. Added Novita AI provider support, thread-safety hardening, improved LLM output parsing robustness, and 20+ bug fixes from community PRs and internal audit.
  • [03/17/2026] v0.3.0MetaClaw Integration — AutoResearchClaw now supports MetaClaw cross-run learning: pipeline failures → structured lessons → reusable skills, injected into all 23 stages. +18.3% robustness in controlled experiments. Opt-in (metaclaw_bridge.enabled: true), fully backward-compatible. See Integration Guide.
  • [03/16/2026] v0.2.0 — Three multi-agent subsystems (CodeAgent, BenchmarkAgent, FigureAgent), hardened Docker sandbox with network-policy-aware execution, 4-round paper quality audit (AI-slop detection, 7-dim review scoring, NeurIPS checklist), and 15+ bug fixes from production runs.
  • [03/15/2026] v0.1.0 — We release AutoResearchClaw: a fully autonomous 23-stage research pipeline that turns a single research idea into a conference-ready paper. No human intervention required.

⚡ One Command. One Paper.

# Fully autonomous — no human intervention
pip install -e . && researchclaw setup && researchclaw init && researchclaw run --topic "Your research idea here" --auto-approve

# Co-Pilot mode — collaborate with AI at key decision points
researchclaw run --topic "Your research idea here" --mode co-pilot

🤔 What Is This?

You think it. AutoResearchClaw writes it. You guide the key decisions.

Drop a research topic — get back a full academic paper with real literature from OpenAlex, Semantic Scholar & arXiv, hardware-aware sandbox experiments (GPU/MPS/CPU auto-detected), statistical analysis, multi-agent peer review, and conference-ready LaTeX targeting NeurIPS/ICML/ICLR. Run it fully autonomous, or use Co-Pilot mode to guide the AI at critical decision points — choose research directions, review experiment designs, and co-write the paper. No hallucinated references.

📄paper_draft.mdFull academic paper (Introduction, Related Work, Method, Experiments, Results, Conclusion)
📐paper.texConference-ready LaTeX (NeurIPS / ICLR / ICML templates)
📚references.bibReal BibTeX references from OpenAlex, Semantic Scholar and arXiv — auto-pruned to match inline citations
🔍verification_report.json4-layer citation integrity + relevance verification (arXiv, CrossRef, DataCite, LLM)
🧪experiment runs/Generated code + sandbox results + structured JSON metrics
📊charts/Auto-generated condition comparison charts with error bars and confidence intervals
📝reviews.mdMulti-agent peer review with methodology-evidence consistency checks
🧬evolution/Self-learning lessons extracted from each run
📦deliverables/All final outputs in one folder — compile-ready for Overleaf

The pipeline runs end-to-end — fully autonomous or with human-in-the-loop collaboration. When experiments fail, it self-heals. When hypotheses don't hold, it pivots. When citations are fake, it kills them. When you want to steer, it pauses and listens.

🌍 Run it anywhere. AutoResearchClaw isn't locked to a single platform. Use it standalone via CLI, plug it into OpenClaw, or wire it up through any ACP-compatible agent — 🤖 Claude Code, 💻 Codex CLI, 🐙 Copilot CLI, ♊ Gemini CLI, 🌙 Kimi CLI, you name it. And because OpenClaw bridges to messaging platforms, you can kick off a full research run from 💬 Discord, ✈️ Telegram, 🐦 Lark (飞书), 💚 WeChat, or wherever your team already hangs out. One topic in, one paper out — no matter where you type it.


🚀 Quick Start

# 1. Clone & install
git clone https://github.com/aiming-lab/AutoResearchClaw.git
cd AutoResearchClaw
python3 -m venv .venv && source .venv/bin/activate
pip install -e .

# 2. Setup (interactive — installs OpenCode beast mode, checks Docker/LaTeX)
researchclaw setup

# 3. Configure
researchclaw init          # Interactive: choose LLM provider, creates config.arc.yaml
# Or manually: cp config.researchclaw.example.yaml config.arc.yaml

# 4. Run
export OPENAI_API_KEY="sk-..."
researchclaw run --config config.arc.yaml --topic "Your research idea" --auto-approve

Output → artifacts/rc-YYYYMMDD-HHMMSS-<hash>/deliverables/ — compile-ready LaTeX, BibTeX, experiment code, charts.

📝 Minimum required config
project:
  name: "my-research"

research:
  topic: "Your research topic here"

llm:
  base_url: "https://api.openai.com/v1"
  api_key_env: "OPENAI_API_KEY"
  primary_model: "gpt-4o"
  fallback_models: ["gpt-4o-mini"]

experiment:
  mode: "sandbox"
  sandbox:
    python_path: ".venv/bin/python"

🧠 What Makes It Different

Capability How It Works
🧑‍✈️ Co-Pilot Mode 6 intervention modes — from fully autonomous to step-by-step. Guide the AI at critical decisions (hypotheses, baselines, paper writing) or let it run free. SmartPause auto-detects when human input would help.
🔄 PIVOT / REFINE Loop Stage 15 autonomously decides: PROCEED, REFINE (tweak params), or PIVOT (new direction). Artifacts auto-versioned.
🤖 Multi-Agent Debate Hypothesis generation, result analysis, and peer review each use structured multi-perspective debate.
🧬 Self-Learning Lessons extracted per run (decision rationale, runtime warnings, metric anomalies) with 30-day time-decay. Future runs learn from past mistakes.
📚 Knowledge Base Every run builds structured KB across 6 categories (decisions, experiments, findings, literature, questions, reviews).
🛡️ Sentinel Watchdog Background quality monitor: NaN/Inf detection, paper-evidence consistency, citation relevance scoring, anti-fabrication guard.
🔍 Claim Verification Inline fact-checking: extracts claims from AI-generated text and cross-references against collected literature. Flags ungrounded citations and fabricated numbers.
🌿 Branch Exploration Fork the pipeline to explore multiple research directions simultaneously, compare results side-by-side, and merge the best path forward.

🦞 OpenClaw Integration

AutoResearchClaw is an OpenClaw-compatible service. Install it in OpenClaw and launch autonomous research with a single message — or use it standalone via CLI, Claude Code, or any AI coding assistant.

🚀 Use with OpenClaw (Recommended)

If you already use OpenClaw as your AI assistant:

1️⃣  Share the GitHub repo URL with OpenClaw
2️⃣  OpenClaw auto-reads RESEARCHCLAW_AGENTS.md → understands the pipeline
3️⃣  Say: "Research [your topic]"
4️⃣  Done — OpenClaw clones, installs, configures, runs, and returns results

That's it. OpenClaw handles git clone, pip install, config setup, and pipeline execution automatically. You just chat.

💡 What happens under the hood
  1. OpenClaw reads RESEARCHCLAW_AGENTS.md → learns the research orchestrator role
  2. OpenClaw reads README.md → understands installation and pipeline structure
  3. OpenClaw copies config.researchclaw.example.yamlconfig.yaml
  4. Asks for your LLM API key (or uses your environment variable)
  5. Runs pip install -e . + researchclaw run --topic "..." --auto-approve
  6. Returns the paper, LaTeX, experiments, and citations

🔌 OpenClaw Bridge (Advanced)

For deeper integration, AutoResearchClaw includes a bridge adapter system with 6 optional capabilities:

# config.arc.yaml
openclaw_bridge:
  use_cron: true              # ⏰ Scheduled research runs
  use_message: true           # 💬 Progress notifications (Discord/Slack/Telegram)
  use_memory: true            # 🧠 Cross-session knowledge persistence
  use_sessions_spawn: true    # 🔀 Spawn parallel sub-sessions for concurrent stages
  use_web_fetch: true         # 🌐 Live web search during literature review
  use_browser: false          # 🖥️ Browser-based paper collection

Each flag activates a typed adapter protocol. When OpenClaw provides these capabilities, the adapters consume them without code changes. See docs/integration-guide.md for full details.

ACP (Agent Client Protocol)

AutoResearchClaw can use any ACP-compatible coding agent as its LLM backend — no API keys required. The agent communicates via acpx, maintaining a single persistent session across all 23 pipeline stages.

Agent Command Notes
Claude Code claude Anthropic
Codex CLI codex OpenAI
Copilot CLI gh GitHub
Gemini CLI gemini Google
OpenCode opencode SST
Kimi CLI kimi Moonshot
# config.yaml — ACP example
llm:
  provider: "acp"
  acp:
    agent: "claude"   # Any ACP-compatible agent CLI command
    cwd: "."          # Working directory for the agent
  # No base_url or api_key needed — the agent handles its own auth.
# Just run — the agent uses its own credentials
researchclaw run --config config.yaml --topic "Your research idea" --auto-approve

🛠️ Other Ways to Run

Method How
Standalone CLI researchclaw run --topic "..." --auto-approve (autonomous) or --mode co-pilot (collaborative)
Python API from researchclaw.pipeline import Runner; Runner(config).run()
Claude Code Reads RESEARCHCLAW_CLAUDE.md — just say "Run research on [topic]"
Copilot CLI researchclaw run --topic "..." with llm.acp.agent: "gh"
OpenCode Reads .claude/skills/ — same natural language interface
Any AI CLI Provide RESEARCHCLAW_AGENTS.md as context → agent auto-bootstraps

🔬 Pipeline: 23 Stages, 8 Phases

Phase A: Research Scoping          Phase E: Experiment Execution
  1. TOPIC_INIT                      12. EXPERIMENT_RUN
  2. PROBLEM_DECOMPOSE               13. ITERATIVE_REFINE  ← self-healing

Phase B: Literature Discovery      Phase F: Analysis & Decision
  3. SEARCH_STRATEGY                 14. RESULT_ANALYSIS    ← multi-agent
  4. LITERATURE_COLLECT  ← real API  15. RESEARCH_DECISION  ← PIVOT/REFINE
  5. LITERATURE_SCREEN   [gate]
  6. KNOWLEDGE_EXTRACT               Phase G: Paper Writing
                                     16. PAPER_OUTLINE
Phase C: Knowledge Synthesis         17. PAPER_DRAFT
  7. SYNTHESIS                       18. PEER_REVIEW        ← evidence check
  8. HYPOTHESIS_GEN    ← debate      19. PAPER_REVISION

Phase D: Experiment Design         Phase H: Finalization
  9. EXPERIMENT_DESIGN   [gate]      20. QUALITY_GATE      [gate]
 10. CODE_GENERATION                 21. KNOWLEDGE_ARCHIVE
 11. RESOURCE_PLANNING               22. EXPORT_PUBLISH     ← LaTeX
                                     23. CITATION_VERIFY    ← relevance check

Gate stages (5, 9, 20) pause for human approval or auto-approve with --auto-approve. On rejection, the pipeline rolls back.

Co-Pilot mode (--mode co-pilot): Deep human-AI collaboration at Stages 7-8 (Idea Workshop), Stage 9 (Baseline Navigator), and Stages 16-17 (Paper Co-Writer). Other stages auto-execute with SmartPause monitoring.

Decision loops: Stage 15 can trigger REFINE (→ Stage 13) or PIVOT (→ Stage 8), with automatic artifact versioning.

📋 What Each Phase Does
Phase What Happens
A: Scoping LLM decomposes the topic into a structured problem tree with research questions
A+: Hardware Auto-detects GPU (NVIDIA CUDA / Apple MPS / CPU-only), warns if local hardware is limited, adapts code generation accordingly
B: Literature Multi-source search (OpenAlex → Semantic Scholar → arXiv) for real papers, screens by relevance, extracts knowledge cards
C: Synthesis Clusters findings, identifies research gaps, generates testable hypotheses via multi-agent debate
D: Design Designs experiment plan, generates hardware-aware runnable Python (GPU tier → package selection), estimates resource needs
E: Execution Runs experiments in sandbox, detects NaN/Inf and runtime bugs, self-heals code via targeted LLM repair
F: Analysis Multi-agent analysis of results; autonomous PROCEED / REFINE / PIVOT decision with rationale
G: Writing Outlines → section-by-section drafting (5,000-6,500 words) → peer reviews (with methodology-evidence consistency) → revises with length guard
H: Finalization Quality gate, knowledge archival, LaTeX export with conference template, citation integrity + relevance verification

✨ Key Features

Feature Description
📚 Multi-Source Literature Real papers from OpenAlex, Semantic Scholar & arXiv — query expansion, deduplication, circuit breaker with graceful degradation
🔍 4-Layer Citation Verification arXiv ID check → CrossRef/DataCite DOI → Semantic Scholar title match → LLM relevance scoring. Hallucinated refs auto-removed.
🖥️ Hardware-Aware Execution Auto-detects GPU (NVIDIA CUDA / Apple MPS / CPU-only) and adapts code generation, imports, and experiment scale accordingly
🦾 OpenCode Beast Mode Complex experiments auto-routed to OpenCode — generates multi-file projects with custom architectures, training loops, and ablation studies. Install via researchclaw setup.
🧪 Sandbox Experiments AST-validated code, immutable harness, NaN/Inf fast-fail, self-healing repair, iterative refinement (up to 10 rounds), partial result capture
📝 Conference-Grade Writing NeurIPS/ICML/ICLR templates, section-by-section drafting (5,000-6,500 words), anti-fabrication guard, revision length guard, anti-disclaimer enforcement
📐 Template Switching neurips_2025, iclr_2026, icml_2026 — Markdown → LaTeX with math, tables, figures, cross-refs, \cite{}
🛡️ Anti-Fabrication VerifiedRegistry enforces ground-truth experiment data in papers. Auto-diagnoses failed experiments and repairs them before writing. Unverified numbers sanitized.
🚦 Quality Gates 3 human-in-the-loop gates (Stages 5, 9, 20) with rollback. Skip with --auto-approve.
🧑‍✈️ HITL Co-Pilot 6 intervention modes with per-stage policies. Idea Workshop, Baseline Navigator, Paper Co-Writer for deep collaboration. SmartPause, cost guardrails, escalation policies, and intervention learning for production safety. CLI/WebSocket/MCP adapters.
💰 Cost Guardrails Budget monitoring with configurable threshold alerts (50%/80%/100%). Pipeline auto-pauses when cost exceeds budget.
🔐 Reproducibility SHA256 checksums for all stage artifacts. Immutable manifests for verification. Multi-level undo with versioned snapshots.

🧑‍✈️ Human-in-the-Loop Co-Pilot

AutoResearchClaw v0.4.0 introduces a complete Human-in-the-Loop (HITL) system that transforms the pipeline from purely autonomous to a human-AI collaborative research engine. Choose your level of involvement:

Intervention Modes

Mode Command What It Does
Full Auto --auto-approve Original behavior — no human intervention
Gate Only --mode gate-only Pause at 3 gate stages (5, 9, 20) for approval
Checkpoint --mode checkpoint Pause at each phase boundary (8 checkpoints)
Co-Pilot --mode co-pilot Deep collaboration at critical stages, auto elsewhere
Step-by-Step --mode step-by-step Pause after every stage — learn the pipeline
Express --mode express Quick review — only 3 most critical gates
Custom --mode custom Define per-stage policies via stage_policies config

Co-Pilot Workflow(updated Apr 13, added experiment to prove the best)

You: researchclaw run --topic "Quantum noise as neural network regularization" --mode co-pilot

Pipeline runs Stages 1-7 automatically...

  ┌─────────────────────────────────────────────────────────────┐
  │  HITL | Stage 08: HYPOTHESIS_GEN                            │
  │  Post-stage review                                          │
  │                                                             │
  │  Hypotheses mentioned: 3                                    │
  │  Novelty score: 0.72 (moderate)                             │
  │                                                             │
  │  [a] Approve  [r] Reject  [e] Edit  [c] Collaborate         │
  │  [i] Inject guidance  [v] View output  [q] Abort            │
  └─────────────────────────────────────────────────────────────┘

You: c  (start collaborative chat)
You: Hypothesis 3 is interesting but needs Dropout/Label Smoothing as baselines
AI:  Updated — added Dropout, Label Smoothing, MixUp, CutMix as baselines...
You: approve

Pipeline continues with your refined hypothesis...

CLI Commands

# Start with HITL mode
researchclaw run --topic "..." --mode co-pilot

# Attach to a paused pipeline (from another terminal)
researchclaw attach artifacts/rc-2026-xxx

# Check pipeline and HITL status
researchclaw status artifacts/rc-2026-xxx

# Approve/reject from another terminal or script
researchclaw approve artifacts/rc-2026-xxx --message "LGTM"
researchclaw reject artifacts/rc-2026-xxx --reason "Missing key baseline"

# Inject guidance for a stage (even before it runs)
researchclaw guide artifacts/rc-2026-xxx --stage 9 --message "Use ResNet-50 as primary baseline"

Key Capabilities

Feature Description
Idea Workshop Brainstorm, evaluate, and refine hypotheses collaboratively (Stage 7-8)
Baseline Navigator AI suggests baselines + human adds/removes + reproducibility checklist (Stage 9)
Paper Co-Writer Section-by-section drafting with human editing and AI polishing (Stage 16-19)
SmartPause Confidence-driven dynamic pausing — auto-detects when human input would help
Claim Verification Inline fact-checking against collected literature — flags ungrounded claims
Cost Guardrails Budget monitoring with 50%/80%/100% threshold alerts
Intervention Learning ALHF — learns from your review patterns to optimize future pause decisions
Branch Exploration Fork pipeline to explore multiple hypotheses, compare, merge the best
Escalation Policy Tiered notification (terminal → Slack → email → auto-halt) when unattended
3 Adapters CLI (terminal), WebSocket (web dashboard), MCP (external agents)

Configuration

# config.arc.yaml
hitl:
  enabled: true
  mode: co-pilot                     # full-auto | gate-only | checkpoint | co-pilot | custom
  cost_budget_usd: 50.0              # Pause when cost exceeds budget (0 = no limit)

  notifications:
    on_pause: true
    on_quality_drop: true
    channels: ["terminal"]            # terminal | slack | webhook

  timeouts:
    default_human_timeout_sec: 86400  # 24h default wait
    auto_proceed_on_timeout: false

  collaboration:
    max_chat_turns: 50
    save_chat_history: true

  # Per-stage custom policies (optional, for 'custom' mode)
  stage_policies:
    8: { require_approval: true, enable_collaboration: true }
    9: { require_approval: true, allow_edit_output: true }

Backward Compatibility

  • Default: OFF. Without hitl.enabled: true or --mode, the pipeline behaves exactly as before.
  • --auto-approve still works. It overrides HITL mode.
  • All 2,699 existing tests pass with HITL code present.

🧠 MetaClaw Integration

AutoResearchClaw + MetaClaw = A pipeline that learns from every run.

MetaClaw adds cross-run knowledge transfer to AutoResearchClaw. When enabled, the pipeline automatically captures lessons from failures and warnings, converts them into reusable skills, and injects those skills into all 23 pipeline stages on subsequent runs — so the same mistakes are never repeated.

How It Works

Run N executes → failures/warnings captured as Lessons
                      ↓
          MetaClaw Lesson → Skill conversion
                      ↓
          arc-* Skill files stored in ~/.metaclaw/skills/
                      ↓
Run N+1 → build_overlay() injects skills into every LLM prompt
                      ↓
          LLM avoids known pitfalls → higher quality, fewer retries

Quick Setup

# 1. Install MetaClaw (if not already)
pip install metaclaw

# 2. Enable in your config
# config.arc.yaml
metaclaw_bridge:
  enabled: true
  proxy_url: "http://localhost:30000"        # MetaClaw proxy (optional)
  skills_dir: "~/.metaclaw/skills"          # Where skills are stored
  fallback_url: "https://api.openai.com/v1" # Direct LLM fallback
  fallback_api_key: ""                      # API key for fallback URL
  lesson_to_skill:
    enabled: true
    min_severity: "warning"                 # Convert warnings + errors
    max_skills_per_run: 3
# 3. Run as usual — MetaClaw works transparently
researchclaw run --config config.arc.yaml --topic "Your idea" --auto-approve

After each run, check ~/.metaclaw/skills/arc-*/SKILL.md to see the skills your pipeline has learned.

Experiment Results

In controlled A/B experiments (same topic, same LLM, same configuration):

Metric Baseline With MetaClaw Improvement
Stage retry rate 10.5% 7.9% -24.8%
Refine cycle count 2.0 1.2 -40.0%
Pipeline stage completion 18/19 19/19 +5.3%
Overall robustness score (composite) 0.714 0.845 +18.3%

Composite robustness score is a weighted average of stage completion rate (40%), retry reduction (30%), and refine cycle efficiency (30%).

Backward Compatibility

  • Default: OFF. If metaclaw_bridge is absent or enabled: false, the pipeline behaves exactly as before.
  • No new dependencies. MetaClaw is optional — the core pipeline works without it.
  • All 2,699 existing tests pass with the integration code present.

🧩 Skills Library

AutoResearchClaw now supports loading open-source and custom skills to further enhance your research experience. We also ship with 20 pre-loaded built-in skills (scientific writing, literature search, chemistry, biology, and more) as ready-to-use references, offering a high degree of flexibility out of the box. Disable any skill by adding enabled: false to its frontmatter.

Sample built-in skills:

Category Skill Description
Writing scientific-writing IMRAD structure, citation formatting, reporting guidelines
Domain chemistry-rdkit Molecular analysis, SMILES, fingerprints, drug discovery
Experiment literature-search Systematic review, PRISMA methodology

See all 20 skills with researchclaw skills list.

Load Your Own Skills

# Option 1: Install a skill (persists across projects)
researchclaw skills install /path/to/my-skill/

# Option 2: Drop a SKILL.md into the project
mkdir -p .claude/skills/my-custom-skill
# Then create a SKILL.md with YAML frontmatter (name, description, trigger-keywords, applicable-stages)

# Option 3: Configure shared skill directories in config.arc.yaml
# skills:
#   custom_dirs:
#     - /path/to/team-shared-skills

Using Skills

Skills are loaded and injected into LLM prompts automatically — no manual activation needed. Use the CLI to inspect:

researchclaw skills list               # Show all loaded skills with sources
researchclaw skills validate ./my-skill # Check SKILL.md format

Browse community skills: K-Dense-AI/claude-scientific-skills (150+ scientific skills across multiple disciplines).


⚙️ Configuration Reference

Click to expand full configuration reference
# === Project ===
project:
  name: "my-research"              # Project identifier
  mode: "docs-first"               # docs-first | semi-auto | full-auto

# === Research ===
research:
  topic: "..."                     # Research topic (required)
  domains: ["ml", "nlp"]           # Research domains for literature search
  daily_paper_count: 8             # Target papers per search query
  quality_threshold: 4.0           # Minimum quality score for papers

# === Runtime ===
runtime:
  timezone: "America/New_York"     # For timestamps
  max_parallel_tasks: 3            # Concurrent experiment limit
  approval_timeout_hours: 12       # Gate stage timeout
  retry_limit: 2                   # Retry count on stage failure

# === LLM ===
llm:
  provider: "openai-compatible"    # openai | openrouter | deepseek | minimax | acp | openai-compatible
  base_url: "https://..."          # API endpoint (required for openai-compatible)
  api_key_env: "OPENAI_API_KEY"    # Env var for API key (required for openai-compatible)
  api_key: ""                      # Or hardcode key here
  primary_model: "gpt-4o"          # Primary model
  fallback_models: ["gpt-4o-mini"] # Fallback chain
  s2_api_key: ""                   # Semantic Scholar API key (optional, higher rate limits)
  acp:                             # Only used when provider: "acp"
    agent: "claude"                # ACP agent CLI command (claude, codex, gemini, etc.)
    cwd: "."                       # Working directory for the agent

# === Literature search ===
literature_search:
  sources: ["openalex", "semantic_scholar", "arxiv"]  # Stage 4 backend order
  max_results_per_query: 40          # Results per query before deduplication
  inter_query_delay_sec: 1.5         # Delay between expanded queries
  openalex_email: "researchclaw@users.noreply.github.com"
  openalex_api_key_env: "OPENALEX_API_KEY"
  openalex_api_key: ""              # Optional; env var is preferred
  s2_api_key_env: "S2_API_KEY"
  s2_api_key: ""                    # Optional; falls back to llm.s2_api_key

# === Experiment ===
experiment:
  mode: "sandbox"                  # simulated | sandbox | docker | ssh_remote
  time_budget_sec: 300             # Max execution time per run (default: 300s)
  max_iterations: 10               # Max optimization iterations
  metric_key: "val_loss"           # Primary metric name
  metric_direction: "minimize"     # minimize | maximize
  sandbox:
    python_path: ".venv/bin/python"
    gpu_required: false
    allowed_imports: [math, random, json, csv, numpy, torch, sklearn]
    max_memory_mb: 4096
  docker:
    image: "researchclaw/experiment:latest"
    network_policy: "setup_only"   # none | setup_only | pip_only | full
    gpu_enabled: true
    memory_limit_mb: 8192
    auto_install_deps: true        # Auto-detect imports → requirements.txt
  ssh_remote:
    host: ""                       # GPU server hostname
    gpu_ids: []                    # Available GPU IDs
    remote_workdir: "/tmp/researchclaw_experiments"
  opencode:                          # OpenCode Beast Mode (auto-installed via `researchclaw setup`)
    enabled: true                    # Master switch (default: true)
    auto: true                       # Auto-trigger without confirmation (default: true)
    complexity_threshold: 0.2        # 0.0-1.0 — higher = only trigger on complex experiments
    model: ""                        # Override model (empty = use llm.primary_model)
    timeout_sec: 600                 # Max seconds for OpenCode generation
    max_retries: 1                   # Retry count on failure
    workspace_cleanup: true          # Remove temp workspace after collection
  code_agent:                        # CodeAgent v2 — multi-phase code generation
    enabled: true                    # Use CodeAgent instead of legacy single-prompt codegen
    architecture_planning: true      # Generate deep implementation blueprint before coding
    sequential_generation: true      # Generate files one-by-one following dependency DAG
    hard_validation: true            # AST-based validation gates (blocks identical ablations, hardcoded metrics)
    hard_validation_max_repairs: 2   # Max repair attempts when validation fails
    exec_fix_max_iterations: 3       # Execution-in-the-loop fix attempts
    exec_fix_timeout_sec: 60         # Timeout per exec-fix attempt
  benchmark_agent:                   # BenchmarkAgent — automated dataset & baseline selection
    enabled: true                    # Enable 4-agent benchmark pipeline (Surveyor→Selector→Acquirer→Validator)
    enable_hf_search: true           # Search HuggingFace Datasets
    enable_web_search: true          # Search Google Scholar for benchmarks
    tier_limit: 2                    # Dataset tier filtering (1=small/cached, 2=medium, 3=large)
    min_benchmarks: 1                # Minimum datasets required
    min_baselines: 2                 # Minimum baseline methods required
  figure_agent:                      # FigureAgent — academic figure generation
    enabled: true                    # Enable 5-agent figure pipeline (Planner→CodeGen→Renderer→Critic→Integrator)
    min_figures: 3                   # Minimum figures to generate
    max_figures: 8                   # Maximum figures
    max_iterations: 3                # Critic-driven refinement iterations
    dpi: 300                         # Output resolution
    strict_mode: false               # Fail pipeline if figure generation fails
  repair:                            # Anti-fabrication experiment repair
    enabled: true                    # Auto-diagnose and repair failed experiments
    max_cycles: 3                    # Repair retry loops
    min_completion_rate: 0.5         # >=50% conditions must complete to proceed
    min_conditions: 2                # At least 2 conditions for valid experiment
    use_opencode: true               # Route repairs through OpenCode Beast Mode

# === Web Search (Optional) ===
web_search:
  enabled: true                      # Enable web-augmented literature search
  tavily_api_key_env: "TAVILY_API_KEY"  # Tavily API key env var (optional)
  enable_scholar: true               # Google Scholar search
  enable_pdf_extraction<

(README truncated)

View on GitHub

Releases and announcements

7 total
  1. AutoResearchClaw v0.5.0v0.5.0May 20, 2026

    ## AutoResearchClaw v0.5.0 ### Highlights - **Multi-Domain Architecture**: Expanded beyond ML to support HEP Physics, Biology, Quantum Computing, and Statistics domains with profile-driven deployment - **ARC-Bench Evaluation Framework**: Standardized benchmark suite with 50+ topics across 5 domains (ML01-ML25, P01-P10, Q01-Q10, B01-B07, S01-S03), rubric-based judging, and baseline adapters for AIDE, Agent Laboratory, and AI-Scientist-v2 - **ColliderAgent Integration**: Full HEP physics simulation pipeline support (MadGraph → Pythia → Delphes) with incremental experiment mode and Stage-12 re-entry - **Biology-Agent Integration**: Metabolic modeling with COBRApy/Biopython skills, FBA simulation, and GSMM validation - **Quantum-Qiskit Skill**: Qiskit-based quantum computing experiment support for quantum topics - **Statistics Domain Agent**: Statistical method design, experiment evaluation, and theory analysis - **Requirements Gate**: LLM capability validation before pipeline execution - **Profile-Driven Deployment**: Interactive CLI for domain profile creation and management - **Incremental Experiment Mode**: Resume experiments at Stage-12 with delta-prompt assembly - **Expanded Te

  2. ## v0.4.0 — Human-in-the-Loop Co-Pilot System AutoResearchClaw is no longer purely autonomous. The new HITL Co-Pilot system transforms the pipeline into a human-AI collaborative research engine. ### Highlights - **6+ Intervention Modes**: `full-auto`, `gate-only`, `checkpoint`, `step-by-step`, `co-pilot`, `custom`, `express` - **Idea Workshop**: Brainstorm and refine hypotheses collaboratively (Stages 7-8) - **Baseline Navigator**: Review and customize experiment designs (Stage 9) - **Paper Co-Writer**: Section-by-section collaborative drafting (Stages 16-19) - **SmartPause**: Confidence-driven dynamic intervention - **ALHF Intervention Learning**: Learns from your review patterns - **Claim Verification**: Inline fact-checking against collected literature - **Cost Guardrails**: Budget monitoring with threshold alerts - **Pipeline Branching**: Fork to explore multiple research directions - **CLI Commands**: `attach`, `status`, `approve`, `reject`, `guide` - **3 Adapters**: CLI, WebSocket, MCP ### New Files - `researchclaw/hitl/` — 34 modules (7,500+ lines) - `tests/test_hitl_*.py` — 9 test files (242 tests) - `docs/HITL_GUIDE.md` — 620-line guide - 3 new builtin skills ### Test

  3. ## What's New ### Cross-Platform Support - **ACP-compatible agent backends**: Claude Code, Codex CLI, Copilot CLI, Gemini CLI, Kimi CLI - **OpenClaw bridge**: messaging platform integration (Discord, Telegram, Lark, WeChat) - **CLI-agent code generation backend**: delegates Stages 10 & 13 to external CLI agents with budget control and timeout management ### Anti-Fabrication System - **VerifiedRegistry**: ground-truth whitelist from experiment results with tolerance matching - **Experiment diagnosis & repair loop**: 13 deficiency categories, auto-repair with best-result selection - **Always-on sanitization**: unverified numbers replaced in paper tables ### Stability & Quality - 100+ bug fixes across 8 deep audit rounds - Modular executor refactoring (10K → 400-line facade) - `--resume` auto-detection for interrupted runs - LLM retry hardening with exponential backoff - Community-reported fixes (macOS M3, math/theoretical topics) ### New Subsystems - Assessor (paper quality scoring + venue recommendation) - Calendar (conference deadline tracking) - Collaboration (multi-user research coordination) - Copilot (interactive steering modes) - Dashboard (real-tim

  4. ## What's New ### OpenCode Beast Mode New "Beast Mode" routes complex code generation to [OpenCode](https://github.com/anomalyco/opencode) with automatic 6-signal complexity scoring and graceful fallback to CodeAgent. ### Universal Cross-Domain Support - Domain detector for 7 research domains (ML, physics, chemistry, economics, math, biology, security) - 25+ domain-specific experiment profiles with tailored datasets, metrics, and evaluation protocols - Domain-aware prompt adapters and Docker images ### Code Searcher Agent GitHub-integrated code search for experiment design reference — query generation, pattern extraction, and result caching. ### Web Integration Layer Web search, crawling, PDF extraction, and Google Scholar support for enhanced literature discovery. ### Community Contributions - **Novita AI provider** — added as built-in LLM provider preset (#80) - **Thread-safety hardening** — all module-level globals protected with locks (#77) - **Figure agent config** — fixed 6 missing config fields (#75) - **Robust LLM output parsing** — 4-strategy JSON extractor, ACP-aware YAML extraction (#69) - **Test collection fix** — safe `skipif` guard for Anthropic tests (#53) ###

  5. ## What's New in v0.3.0 ### MetaClaw Cross-Run Learning Integration - New `researchclaw/metaclaw_bridge/` module: skill injection, lesson-to-skill conversion, PRM quality gates, session lifecycle management - Pipeline failures → structured lessons → reusable skills, injected into all 23 stages - **+18.3%** pipeline robustness in controlled experiments - Opt-in via `metaclaw_bridge.enabled: true`, fully backward-compatible ### CodeAgent v2 — Enhanced Code Generation - Enhanced Blueprint: deep implementation specs with per-file pseudocode, tensor shapes, generation order - Sequential File Generation: dependency-ordered with AST-based CodeMem - Hard Validation Gates: block identical ablations, hardcoded metrics, cross-file import errors - Targeted Error Repair: parse traceback to fix surgically instead of full regeneration ### BenchmarkAgent & FigureAgent Improvements - BenchmarkAgent: domain-aware benchmarks, import validation, pretrained resize - FigureAgent: LLM output type safety, Paul Tol colorblind-safe palette, heatmap/ablation chart types - `visualize.py` full rewrite: academic styling, 300 DPI, 6 enhanced chart types ### 50+ Pipeline Bug Fixes (BUG-06 through BUG-51) - Me

Code frequency

additions and deletions
+140K-140KWeek of 2026-03-08: +2 linesWeek of 2026-03-08: -2 linesWeek of 2026-03-15: +139,983 linesWeek of 2026-03-15: -17,700 linesWeek of 2026-03-22: +23,468 linesWeek of 2026-03-22: -631 linesWeek of 2026-03-29: +18,390 linesWeek of 2026-03-29: -537 linesWeek of 2026-04-05: +1,482 linesWeek of 2026-04-05: -161 linesWeek of 2026-04-12: +0 linesWeek of 2026-04-12: -0 linesWeek of 2026-04-19: +491 linesWeek of 2026-04-19: -35 linesWeek of 2026-04-26: +0 linesWeek of 2026-04-26: -0 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +42,725 linesWeek of 2026-05-17: -8,677 linesWeek of 2026-05-24: +746 linesWeek of 2026-05-24: -34 linesWeek of 2026-05-31: +1,033 linesWeek of 2026-05-31: -30 linesWeek of 2026-06-07: +0 linesWeek of 2026-06-07: -0 linesWeek of 2026-06-14: +250 linesWeek of 2026-06-14: -0 linesWeek of 2026-06-21: +0 linesWeek of 2026-06-21: -0 linesWeek of 2026-06-28: +0 linesWeek of 2026-06-28: -0 linesWeek of 2026-07-05: +270 linesWeek of 2026-07-05: -7 linesWeek of 2026-07-12: +371 linesWeek of 2026-07-12: -61 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesMar 8, 2026Jul 26, 2026
+229.2K lines added, -27.9K removed over the last year.

Commits per week

last 52 weeks
1390Week of 2025-08-09: 0 commitsWeek of 2025-08-16: 0 commitsWeek of 2025-08-23: 0 commitsWeek of 2025-08-30: 0 commitsWeek of 2025-09-06: 0 commitsWeek of 2025-09-13: 0 commitsWeek of 2025-09-20: 0 commitsWeek of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 2 commitsWeek of 2026-03-15: 139 commitsWeek of 2026-03-22: 12 commitsWeek of 2026-03-29: 20 commitsWeek of 2026-04-05: 15 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 5 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 6 commitsWeek of 2026-05-24: 10 commitsWeek of 2026-05-31: 5 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 2 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 0 commitsWeek of 2026-07-05: 1 commitsWeek of 2026-07-12: 2 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsAug 9, 2025Aug 2, 2026
219 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 3 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 3 commitsSun 10:00 — 4 commitsSun 11:00 — 3 commitsSun 12:00 — 1 commitsSun 13:00 — 2 commitsSun 14:00 — 0 commitsSun 15:00 — 2 commitsSun 16:00 — 5 commitsSun 17:00 — 3 commitsSun 18:00 — 4 commitsSun 19:00 — 2 commitsSun 20:00 — 4 commitsSun 21:00 — 3 commitsSun 22:00 — 1 commitsSun 23:00 — 2 commitsMon 0:00 — 3 commitsMon 1:00 — 1 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 2 commitsMon 9:00 — 2 commitsMon 10:00 — 3 commitsMon 11:00 — 4 commitsMon 12:00 — 5 commitsMon 13:00 — 3 commitsMon 14:00 — 4 commitsMon 15:00 — 0 commitsMon 16:00 — 4 commitsMon 17:00 — 6 commitsMon 18:00 — 2 commitsMon 19:00 — 3 commitsMon 20:00 — 2 commitsMon 21:00 — 0 commitsMon 22:00 — 1 commitsMon 23:00 — 4 commitsTue 0:00 — 1 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 1 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 3 commitsTue 9:00 — 14 commitsTue 10:00 — 4 commitsTue 11:00 — 5 commitsTue 12:00 — 4 commitsTue 13:00 — 2 commitsTue 14:00 — 0 commitsTue 15:00 — 0 commitsTue 16:00 — 1 commitsTue 17:00 — 2 commitsTue 18:00 — 0 commitsTue 19:00 — 1 commitsTue 20:00 — 0 commitsTue 21:00 — 2 commitsTue 22:00 — 0 commitsTue 23:00 — 1 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 1 commitsWed 3:00 — 2 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 1 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 3 commitsWed 11:00 — 1 commitsWed 12:00 — 2 commitsWed 13:00 — 3 commitsWed 14:00 — 3 commitsWed 15:00 — 1 commitsWed 16:00 — 2 commitsWed 17:00 — 1 commitsWed 18:00 — 0 commitsWed 19:00 — 1 commitsWed 20:00 — 2 commitsWed 21:00 — 4 commitsWed 22:00 — 7 commitsWed 23:00 — 1 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 2 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 0 commitsThu 9:00 — 6 commitsThu 10:00 — 1 commitsThu 11:00 — 2 commitsThu 12:00 — 4 commitsThu 13:00 — 0 commitsThu 14:00 — 2 commitsThu 15:00 — 0 commitsThu 16:00 — 2 commitsThu 17:00 — 1 commitsThu 18:00 — 1 commitsThu 19:00 — 0 commitsThu 20:00 — 2 commitsThu 21:00 — 0 commitsThu 22:00 — 0 commitsThu 23:00 — 1 commitsFri 0:00 — 3 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 1 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 4 commitsFri 10:00 — 1 commitsFri 11:00 — 3 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 2 commitsFri 16:00 — 1 commitsFri 17:00 — 0 commitsFri 18:00 — 0 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 1 commitsFri 22:00 — 2 commitsFri 23:00 — 0 commitsSat 0:00 — 2 commitsSat 1:00 — 0 commitsSat 2:00 — 1 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 1 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 2 commitsSat 22:00 — 0 commitsSat 23:00 — 1 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Mar 18, 2026daily#7+397
Mar 17, 2026daily#10+315
Mar 16, 2026daily#20+118
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    385.5K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python