Imbad0202/academic-research-skillsPublic

Academic Research Skills for Claude Code: research → write → review → revise → finalize

AI summary: A comprehensive suite of Claude Code skills guiding academic research from literature review to publication.

Stars
50.3K
+288 today
Forks
3.9K
Watchers
132
Open issues
32
Open PRs
3
Contributors
~20
Commits
797
Branches
7

PythonOtherCreated Feb 26, 2026Last push 1d agoLatest release v3.22.2+856 stars this week+4.3K this month

Quick answers

What is academic-research-skills?
A comprehensive suite of Claude Code skills guiding academic research from literature review to publication.
What does academic-research-skills do?
This project provides a comprehensive, highly specialized suite of Claude Code skills specifically engineered for the rigorous academic research pipeline. It acts as an intelligent AI copilot designed to handle tedious logistical tasks—such as formatting complex citations, verifying data points, and ensuring logical consistency—while keeping the researcher firmly in control. By enforcing a human-in-the-loop methodology via Socratic dialogue, it ensures researchers actively define their questions and interpret data rather than fully automating the writing process. Additionally, it features advanced Style Calibration to match the user's authentic voice and a Writing Quality Check to eliminate typical AI-generated prose patterns.
Who is academic-research-skills for?
This suite is built for academic researchers, graduate students, and scientists who want a powerful AI copilot to assist with writing logistics. It helps them efficiently streamline technical aspects without compromising their academic integrity or personal authorial voice.
How do I get started with academic-research-skills?
/plugin marketplace add Imbad0202/academic-research-skills
How popular is academic-research-skills on GitHub?
Imbad0202/academic-research-skills has 50,291 stars and 3,878 forks on GitHub, and gained 856 stars in the last 7 days.
What license does academic-research-skills use?
Imbad0202/academic-research-skills is released under the Other license.

Star history

since Jul 28, 2026
020K40KJul 2026Aug 2026Sep 2026Oct 2026
50.3K stars as of Oct 3, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 6 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 4 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 5 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 12 commits2026-03-09: 16 commits2026-03-10: 5 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 6 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 7 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 1 commit2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 6 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 2 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 31 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 1 commit2026-04-14: 1 commit2026-04-15: 49 commits2026-04-16: 1 commit2026-04-17: 0 commits2026-04-18: 1 commit2026-04-19: 1 commit2026-04-20: 27 commits2026-04-21: 2 commits2026-04-22: 22 commits2026-04-23: 76 commits2026-04-24: 0 commits2026-04-25: 25 commits2026-04-26: 1 commit2026-04-27: 4 commits2026-04-28: 0 commits2026-04-29: 1 commit2026-04-30: 2 commits2026-05-01: 3 commits2026-05-02: 5 commits2026-05-03: 2 commits2026-05-04: 0 commits2026-05-05: 12 commits2026-05-06: 1 commit2026-05-07: 2 commits2026-05-08: 2 commits2026-05-09: 2 commits2026-05-10: 1 commit2026-05-11: 7 commits2026-05-12: 6 commits2026-05-13: 0 commits2026-05-14: 3 commits2026-05-15: 5 commits2026-05-16: 1 commit2026-05-17: 4 commits2026-05-18: 4 commits2026-05-19: 8 commits2026-05-20: 8 commits2026-05-21: 4 commits2026-05-22: 0 commits2026-05-23: 5 commits2026-05-24: 5 commits2026-05-25: 2 commits2026-05-26: 0 commits2026-05-27: 4 commits2026-05-28: 11 commits2026-05-29: 3 commits2026-05-30: 2 commits2026-05-31: 5 commits2026-06-01: 5 commits2026-06-02: 5 commits2026-06-03: 1 commit2026-06-04: 9 commits2026-06-05: 9 commits2026-06-06: 9 commits2026-06-07: 8 commits2026-06-08: 10 commits2026-06-09: 5 commits2026-06-10: 14 commits2026-06-11: 5 commits2026-06-12: 2 commits2026-06-13: 4 commits2026-06-14: 0 commits2026-06-15: 5 commits2026-06-16: 2 commits2026-06-17: 1 commit2026-06-18: 3 commits2026-06-19: 0 commits2026-06-20: 2 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 3 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 2 commits2026-07-02: 5 commits2026-07-03: 2 commits2026-07-04: 3 commits2026-07-05: 0 commits2026-07-06: 1 commit2026-07-07: 0 commits2026-07-08: 1 commit2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 6 commits2026-07-12: 4 commits2026-07-13: 4 commits2026-07-14: 0 commits2026-07-15: 3 commits2026-07-16: 3 commits2026-07-17: 0 commits2026-07-18: 11 commits2026-07-19: 0 commits2026-07-20: 4 commits2026-07-21: 0 commits2026-07-22: 3 commits2026-07-23: 0 commits2026-07-24: 4 commits2026-07-25: 3 commits2026-07-26: 5 commits2026-07-27: 5 commits2026-07-28: 2 commits2026-07-29: 2 commits2026-07-30: 1 commit2026-07-31: 0 commits2026-08-01: 4 commits2026-08-02: 4 commits2026-08-03: 3 commits2026-08-04: 3 commits2026-08-05: 3 commits2026-08-06: 8 commits2026-08-07: 2 commits2026-08-08: 5 commits2026-08-09: 10 commits2026-08-10: 6 commits2026-08-11: 9 commits2026-08-12: 0 commits2026-08-13: 7 commits2026-08-14: 10 commits2026-08-15: 2 commits2026-08-16: 0 commits2026-08-17: 9 commits2026-08-18: 6 commits2026-08-19: 3 commits2026-08-20: 4 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 3 commits2026-08-25: 0 commits2026-08-26: 1 commit2026-08-27: 3 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 1 commit2026-08-31: 1 commit2026-09-01: 1 commit2026-09-02: 1 commit2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 3 commits2026-09-07: 0 commits2026-09-08: 4 commits2026-09-09: 2 commits2026-09-10: 0 commits2026-09-11: 1 commit2026-09-12: 2 commits2026-09-13: 5 commits2026-09-14: 4 commits2026-09-15: 2 commits2026-09-16: 4 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 1 commit2026-09-20: 0 commits2026-09-21: 1 commit2026-09-22: 0 commits2026-09-23: 4 commits2026-09-24: 9 commits2026-09-25: 3 commits2026-09-26: 0 commits
762 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Landmark project

    50,291 stars

  • Very active

    762 commits in 52 weeks

  • Well documented

    High community health score

  • Continuous integration

    Automated checks passing

  • Repeat trending

    24 trending appearances

What academic-research-skills does

This project provides a comprehensive, highly specialized suite of Claude Code skills specifically engineered for the rigorous academic research pipeline. It acts as an intelligent AI copilot designed to handle tedious logistical tasks—such as formatting complex citations, verifying data points, and ensuring logical consistency—while keeping the researcher firmly in control. By enforcing a human-in-the-loop methodology via Socratic dialogue, it ensures researchers actively define their questions and interpret data rather than fully automating the writing process. Additionally, it features advanced Style Calibration to match the user's authentic voice and a Writing Quality Check to eliminate typical AI-generated prose patterns.

This suite is built for academic researchers, graduate students, and scientists who want a powerful AI copilot to assist with writing logistics. It helps them efficiently streamline technical aspects without compromising their academic integrity or personal authorial voice.

  • Full pipeline automation: Actively assists in the entire academic workflow from initial literature review to final peer review preparation.
  • Socratic dialogue planning: Provides the /ars-plan command to help researchers structure their paper via interactive, interrogative dialogue.
  • Dynamic style calibration: Analyzes and learns the user's specific writing voice from past work to ensure generated text remains authentic.
  • Advanced quality checking: Scans drafts to explicitly identify and flag predictable prose patterns that feel overly machine-generated.
  • Automated data verification: Handles the exhaustive grunt work of verifying data points and checking logical consistency throughout the manuscript.

Where teams use it

Literature review analysis

Researchers use the tool to efficiently gather, logically organize, and correctly format references and citations automatically.

Manuscript structuring

Authors leverage the specialized /ars-plan command to logically map out their arguments and paper structure prior to drafting.

Authorial voice refinement

Writers apply the Style Calibration feature to maintain a consistent, human-like, and highly personalized voice across their dissertation.

Pre-submission quality assurance

Academics run the rigorous Writing Quality Check before submission to ensure their prose meets high, non-robotic editorial standards.

Getting started: /plugin marketplace add Imbad0202/academic-research-skills

README

main branch

Academic Research Skills for Claude Code

Version DOI License: CC BY-NC 4.0 Sponsor

简体中文版 | 繁體中文版 | 日本語版 | 한국어 | Español

A comprehensive suite of Claude Code skills for academic research, covering the full pipeline from research to publication.

Install in 30 seconds (Claude Code CLI / VS Code / JetBrains, v3.7.0+):

/plugin marketplace add Imbad0202/academic-research-skills
/plugin install academic-research-skills

Then try /ars-plan to walk through your paper structure via Socratic dialogue, or jump to Quick install for prerequisites and the traditional symlink flow.

AI is your copilot, not the pilot. It can draft text, including a whole paper in full mode, but the decisions stay yours, and the pipeline stops for your confirmation at every stage. It handles the grunt work (hunting down references, formatting citations, verifying data, checking logical consistency) so you can focus on the parts that actually require your brain: defining the question, choosing the method, interpreting what the data means, and deciding what comes after "I argue that." You remain the author, and you answer for every claim you submit.

Unlike a humanizer, this tool doesn't help you hide the fact that you used AI. It helps you write better. Style Calibration learns your voice from past work. Writing Quality Check catches the patterns that make prose feel machine-generated. The goal is quality, not cheating.

Why human-in-the-loop, not full automation?

Lu et al. (2026, Nature 651:914-919) built The AI Scientist — the first fully autonomous AI research system to publish a paper through blind peer review at a top-tier ML venue (ICLR 2025 workshop, score 6.33/10 vs workshop average 4.87). Their Limitations section enumerates the failure modes that any fully-autonomous AI research pipeline inherits: implementation bugs, hallucinated results, shortcut reliance, bug-as-insight reframing, methodology fabrication, frame-lock, citation hallucinations.

ARS is built on the premise that a human researcher augmented by AI avoids these failure modes better than either alone. Stage 2.5 and Stage 4.5 integrity gates run a 7-mode blocking checklist (see academic-pipeline/references/ai_research_failure_modes.md); the reviewer offers an opt-in calibration mode that measures its own FNR/FPR against a user-supplied gold set.

Zhao et al. (2026-05) audited 111M references across 2.5M papers on arXiv, bioRxiv, SSRN, and PMC. Their conservative estimate is 146,932 hallucinated citations for 2025 alone, with an observed mid-2024 inflection; for the bioRxiv-to-PMC pairing they report 85.3% preprint-to-published persistence. The paper describes "real citations deployed to support claims the cited references do not actually make" as an open challenge. ARS v3.7.1 added trust-chain frontmatter for source provenance; v3.7.3 added locator infrastructure (three-layer citation anchors) for future claim-level audits and surfaces advisory risk signals at cite time (ARS labels the claim-faithfulness gap internally as "L3"; this is ARS terminology, not the paper's). v3.7.x is motivated by Zhao et al.'s corpus-scale findings; corpus-scale evaluation of ARS itself remains future work.

v3.8 closes the second half of the L3 gap. v3.7.3 made every citation carry a locator anchor; v3.8 adds an opt-in audit pass (ARS_CLAIM_AUDIT=1) that fetches the cited source against each anchor and judges whether the claim is actually supported. Five new HIGH-WARN classes (claim-not-supported, negative-constraint-violation, fabricated-reference, anchorless, constraint-violation-uncited) gate-refuse output through the formatter terminal hard gate. A calibration runner ships with a 25-tuple synthetic gold set and FNR<0.15 + FPR<0.10 acceptance thresholds. Its shipped test drives the runner with a stub judge that returns the gold labels, so it checks the tooling, not a live judge; no live-judge calibration result is recorded yet, and the ramp-on plan waits for one (v3.8 spec §5).

Ren et al. (2026, Self-Improvements in Modern Agentic Systems: A Survey) supplies a third, survey-level anchor. Its scientific-discovery synthesis (§7.4) concludes that discovery agents cannot easily verify novelty, correctness, or reproducibility on their own and may exploit weak proxies instead, must manage evidence across heterogeneous tools and literature, and raise governance issues — "scientific writing can also amplify misinformation when the evidence is weak." Its generation-loop chapters (§5.1–§5.2) list human auditing and retained human anchors among the practical safeguards for self-generated evaluation loops, and its historical chapter (§2.2) records the oldest form of the same lesson: the practical success of Lenat's EURISKO depended heavily on the user serving as the external evaluation signal, pruning unproductive heuristic drift — a limitation the survey notes persists in modern agentic systems. ARS cites the survey as design rationale for its human-in-the-loop stance, not as empirical proof that human-in-the-loop pipelines outperform autonomous ones; the survey's actionable deltas for ARS are tracked in #539–#541 and #547–#550.

Gartenberg et al. (2026, Organization Science 37(3):795-812, "More versus better") supplies a fourth anchor, and the first from the journal side. The Organization Science AI Task Force scored every first submission (6,957) and every text-format review (10,389) the journal received between January 2021 and February 2026 with a commercial AI-writing classifier and standard readability indices. Manuscripts scored as heavily AI-written read worse on those indices and were desk-rejected more often; reviews scored as more AI-written leaned toward theory and away from data; and the editors conclude that current AI tools, amplified by publish-or-perish incentives, "appear to be pushing the system toward an equilibrium of more rather than better research." Their §5 contrasts "cognitive surrender" (Shaw & Nave, 2026, as cited there) with human-first use and asks authors to disclose how a manuscript was produced. The evidence is observational, aggregate, and from one journal, and the classifier is a proprietary instrument. ARS cites the editorial as design rationale for recording volume as a non-goal (see POSITIONING.md) and for the Collaboration Depth Observer and the claim-strength ladder, not as evidence about ARS output; the actionable deltas are tracked in #829–#833.

Wang, Li et al. (2026-09, The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing, a survey of 230 sources) supplies a fifth anchor, and the first that treats research production and peer review as one coupled system. Its evaluative-authority ladder (§4.1) runs from author-facing feedback through reviewer assistance and official AI reviews to scoring and decision support, with the survey's point that capability at one rung does not justify use at the next; ARS's simulated panel sits on the lowest rung by design (see POSITIONING.md). Two of its findings shape the reviewer roadmap. First, as the survey summarizes Dycke & Gurevych (2026, §4.5), 391 edits that break a paper's scientific support relations produced no statistically significant difference in the tested automated reviewers' aspects, sentiment, or scores compared with 540 soundness-neutral controls, while presentation-only rewrites with the science held fixed moved AI-review scores (§5.2); the survey's §9.2 conclusion is that a static evaluation can overstate an AI reviewer's reliability once authors can observe and adapt to it. ARS tracks the corresponding Round-1 paired controls in #871 and the author-identity-cue controls (§7.2) in #872; both are measurements, not new mechanisms. Second, its §9.1 cites Brodeur et al. (2026, PNAS 123(22):e2524747123), a randomized study in which 288 researchers in 103 teams reproduced published quantitative social-science results under three conditions: human-only, AI-assisted (ChatGPT as a collaborative tool), and AI-led (ChatGPT with minimal human oversight). Human-only and AI-assisted teams reproduced 94% and 91%, AI-led teams 37%, and the AI-assisted teams detected fewer major coding errors than the human-only teams. In that study, then, AI assistance did not verify better than humans alone and AI-led verification did much worse; ARS reads this as a reason to keep verification human-led at every checkpoint, not as evidence that its own checkpoints or integrity gates are effective. The survey is a synthesis rather than an experiment, its structured search stops at 2026-07-01 and the later targeted update did not rerun every query (§10), its deployment evidence is concentrated in a small number of AI/CS conferences, OpenReview-based settings, and selected journals, and its "arms race" framing is a lens, not a finding; ARS cites it as design rationale, not as evidence about ARS output.

v3.3 was inspired by PaperOrchestra (Song, Song, Pfister & Yoon, 2026, Google): Semantic Scholar API verification, anti-leakage protocol, VLM figure verification, and revision-trajectory tracking. ARS now implements that last idea through categorical, evidence-anchored criterion trajectories rather than score deltas.


Architecture & pipeline

👉 docs/ARCHITECTURE.md — the full pipeline view: flow diagram, stage-by-stage matrix, data-access flow, skill dependency graph, quality gates, and mode list.

The architecture doc supersedes the sprawling pipeline description that used to live here. Everything about what runs in which stage now lives in one place.

Quick install

Prerequisites

  • Claude Code (latest; plugin packaging requires recent versions)
  • ANTHROPIC_API_KEY exported, or set on first claude run
  • Optional: Pandoc for DOCX, tectonic + Source Han Serif TC for APA 7.0 PDF (Markdown output works without either)
  • Optional (real Python): needed only for the write-scope guard and a few opt-in commands; the core skills are prompt-driven. Details, including the Windows notes on Git Bash and the Microsoft Store Python stub, are in docs/SETUP.md § Python (optional).

Which controls are active in your install channel? Availability varies by install channel. See the per-channel map: docs/CONTROL_AVAILABILITY.md.

Plugin install (v3.7.0+, recommended):

/plugin marketplace add Imbad0202/academic-research-skills
/plugin install academic-research-skills

Verify it works: run /ars-plan and describe a paper you're working on — ARS will start a Socratic dialogue to map out chapter structure. For a single-shot test instead, try /ars-lit-review "your topic".

👉 docs/SETUP.md — full guide: install Claude Code, set up API keys, optional Pandoc/tectonic for DOCX/PDF, cross-model verification (ARS_CROSS_MODEL), and six installation methods (Plugin, project skills, global skills, claude.ai Project, repo-cloned, Claude Science import).

👉 docs/DATA_FLOWS.md — what leaves your machine (bibliographic resolvers, optional consent-gated cross-model calls, the plugin update check), what is cached locally, for how long, and how to turn each path off.

👉 docs/RISK_REGISTER.md — the standing risks the suite knows about, which existing controls address each one, the evidence status behind those controls, and what remains open.

Using Claude Science? The four skills import directly: Skills → Import from GitHub, paste https://github.com/Imbad0202/academic-research-skills, Preview, then Import 4 skills (requires v3.14.0+ of this repo — the importer reads the explicit skill paths in the marketplace manifest). Imports are point-in-time snapshots: re-import after ARS updates. Imported skills carry the ARS methodology (research / writing / review protocols); Claude Code-specific machinery — slash commands, hooks, subagent orchestration — does not transfer. See docs/SETUP.md Method 5 for details.

Using Pi? Install the in-tree, community-maintained wrapper with pi install git:github.com/Imbad0202/academic-research-skills. It keeps the original ARS content authoritative and documents Pi-specific orchestration and hook limitations. See pi/README.md.

Using Codex CLI? Install the sibling distribution instead: Imbad0202/academic-research-skills-codex — same workflow content, Codex-native packaging as a single $academic-research-suite skill with ars-* aliases.

Third-party platforms and integrations that wrap or host ARS are listed in THIRD_PARTY.md — community-submitted and not reviewed or endorsed by the maintainer.

Governance: who decides, what cross-model review does and does not provide, and the project's end-of-life posture are stated in GOVERNANCE.md; security reporting and triage in SECURITY.md.

Performance & cost

👉 docs/PERFORMANCE.md — per-mode token budgets, full-pipeline estimate (about US$3–7 for a 15k-word paper at 2026-09 list prices, before cache discounts), and recommended Claude Code settings (Auto mode; Agent Team optional).

Guides & articles


Features at a glance

  • Deep Research — 13-agent research team with Socratic guided mode, PRISMA systematic review, intent detection, dialogue health monitoring, optional cross-model DA, Semantic Scholar API verification.
  • Academic Paper — 12-agent paper writing with Style Calibration, Writing Quality Check, LaTeX hardening, visualization, revision coaching, citation conversion, anti-leakage protocol, and VLM figure verification.
  • Academic Paper Reviewer — 7-agent multi-perspective peer review with criterion-bound, evidence-anchored narrative judgements (Journal-Fit Reviewer + 3 dynamic reviewers + Devil's Advocate), concession threshold protocol, attack intensity preservation, optional cross-model DA critique / calibration, R&R traceability matrix, read-only constraint. Current live reviews remain NOT_CALIBRATED; full calibration produces a bounded candidate profile, while live-profile application is not yet wired.
  • Academic Pipeline — 10-stage pipeline orchestrator with adaptive checkpoints, claim verification, Material Passport, optional repro_lock, optional cross-model integrity verification, mid-conversation reinforcement, and narrative criterion-by-criterion regression checks (the typed trajectory carrier is deferred).
  • Data Access Level Metadata (v3.3.2+) — every skill declares data_access_level (raw / redacted / verified_only); enforced by scripts/check_data_access_level.py. Pattern adapted from Anthropic's automated-w2s-researcher (2026). See shared/ground_truth_isolation_pattern.md.
  • Task Type Annotation (v3.3.2+) — every skill declares task_type (open-ended or outcome-gradable). All current ARS skills are open-ended.
  • Benchmark Report Schema (v3.3.5+) — JSON Schema + lint for honest benchmark comparisons. See shared/benchmark_report_pattern.md.
  • Artifact Reproducibility Lockfile (v3.3.5+) — optional repro_lock sub-block on Material Passport. Configuration documentation, not replay guarantee — LLM outputs are not byte-reproducible. See shared/artifact_reproducibility_pattern.md.
  • Model Tiering (#517, v3.16+) — optional ARS_MODEL_TIERING switch with two directions: economy (execution-type agents dispatch one tier below the session model, floor Opus-class) and quality-boost (judgment-type agents at integrity gates and final review step up to the frontier tier). Default unset = byte-equivalent to pre-#517 behavior. See shared/model_tiering.md.
  • Canonical Cross-Model Handoff Envelope (#527, v3.17+) — the owner→dispatcher→owner blind-checkpoint transport path (#523) now has a machine-stable [CROSS-MODEL-HANDOFF v1] envelope with a normative Python grammar (scripts/cross_model_handoff.py) instead of prose-only enforcement, pinning agreement/divergence/malformed-result routing across all three checkpoint owners. See shared/cross_model_verification.md §"Cross-model handoff envelope".
  • Experiment Provenance Intake (#260) — optional experiment_provenance[] on the Material Passport records experiments the scholar ran externally (ARS never runs experiments), and manuscript claims join to them via claim_intent_manifest.planned_experiment_ids[]. The integrity gate (Stage 2.5/4.5) audits each experiment-backed claim against declared provenance — ALIGNED / OVERSTATED / NOT_SUPPORTED_BY_PROVENANCE / PROVENANCE_INSUFFICIENT — without judging whether the experiment itself was correct. A fail-closed experiment_intake_declaration makes "did you run experiments?" an explicit Stage 1 decision (even literature-only runs declare no_experiments_declared). See shared/handoff_schemas.md §"Experiment Provenance Intake (#260)".

Integrity and verification boundary: ARS checks the manuscript and the reported process—including citation existence, claim–source alignment, reported methodology, declared experiment–result alignment, figure/table fidelity, and reporting/process/package conformance. Some checks are sampled or LLM-mediated. ARS does not establish that procedures were actually performed, raw data are authentic, or results reproduce; a consistently reported fabrication can pass these checks. See POSITIONING.md § Integrity checks and the empirical-work boundary.


Showcase: real pipeline output

See the complete artifacts from a real pipeline run, including peer review reports, integrity verification reports, and the final paper:

A March 2026 record, not current performance. This run (2026-03-07 to 03-08) used academic-pipeline v2.3, before ARS v3.3 added the Semantic Scholar check and v3.11 added the deterministic four-index citation gate. Its numbers describe that version; the current gates have not been measured on this paper. The byline names Claude (Anthropic) as author because the researcher asked for that during the experiment. The paper is not an Anthropic publication, and ARS's positioning is that the tool does not replace the researcher and does not claim authorship (POSITIONING.md).

Browse all pipeline artifacts →

Artifact Description
Final Paper (EN) APA 7.0 formatted, LaTeX-compiled
Final Paper (ZH) Chinese version, APA 7.0
Integrity Report — Pre-Review Stage 2.5: flagged 15 problem references (8 with bibliographic errors, 6–8 likely fabricated) + 3 statistical errors
Integrity Report — Final Stage 4.5: zero regressions confirmed
Peer Review Round 1 Journal-Fit Reviewer + 3 Reviewers + Devil's Advocate
Re-Review Verification after revisions
Peer Review Round 2 Follow-up review
Response to Reviewers Point-by-point author response
Post-Publication Audit Report Full-reference audit run separately with Claude Code + WebSearch: 21 of the 68 final references still had problems after 3 rounds of integrity checks

Companion: Experiment Agent

If your research involves running experiments (code or human studies) before writing, the Experiment Agent skill fills the gap between ARS Stage 1 (RESEARCH) and Stage 2 (WRITE).

ARS Stage 1 RESEARCH  →  RQ Brief + Methodology Blueprint
        ↓
  experiment-agent     →  run/manage experiments → validate results
        ↓
ARS Stage 2 WRITE     →  write paper with verified experiment results

What it does: executes code experiments (Python, R, etc.) with real-time monitoring, manages human study protocols with IRB ethics checklist, interprets statistics with 11-type fallacy detection, and verifies reproducibility.

How to use together: pause the ARS pipeline after Stage 1, run experiments in a separate experiment-agent session, then bring the results (with Material Passport) back to ARS Stage 2. ARS requires zero modification. See the experiment-agent README for setup instructions.

Stage 1 intake declaration (#260): at Stage 1, ARS detects whether the run will carry experiment-backed claims and sets a fail-closed experiment_intake_declaration on the Material Passport. If you ran experiments externally, the scholar enters one experiment_provenance[] entry per experiment (experiment_id, nested repro_lock, planned_vs_executed[], negative_results[], known_limitations[]) and the declaration is set to experiments_declared; if not, it is set to no_experiments_declared. The declaration is required on every post-#260 passport — a run that touches no experiments still declares no_experiments_declared, so the integrity gate can never be silently bypassed by a forgotten provenance block. The experiment_ids are frozen at this intake point; the writers later reference them via planned_experiment_ids[].

Teaching-side companion: Teaching Skills applies the ARS architecture (skill ensembles, shared contracts, staged gates, a Course Passport) to the teaching side of academic life — course design → lessons → assessment → delivery → reflection; its sotl mode hands classroom-inquiry projects off to ARS deep-research / academic-paper for the publication phase.


Usage

Quick Start

# Start a full research pipeline
You: "I want to write a research paper on AI's impact on higher education QA"

# Start with Socratic guidance
You: "Guide my research on AI in educational evaluation"

# Write a paper with guided planning
You: "Guide me through writing a paper on demographic decline"

# Review an existing paper
You: "Review this paper" (then provide the paper)

# Check pipeline status
You: "status"

Individual Skills

Deep Research (8 modes)
"Research the impact of AI on higher education"       → full mode
"Give me a quick brief on X"                          → quick mode
"Do a systematic review on X with PRISMA"             → systematic-review mode
"Guide my research on X"                              → socratic mode (guided)
"Fact-check these claims"                             → fact-check mode
"Do a literature review on X"                         → lit-review mode
"Compare these papers in WHY/HOW/WHAT format"         → three-way-scan mode
"Review this paper's research quality"                → review mode
Academic Paper (11 modes)
"Write a paper on X"                                  → full mode
"Guide me through writing a paper"                    → plan mode (guided)
"Build a paper outline"                               → outline-only mode
"I have a draft, here are reviewer comments"          → revision mode
"Parse these reviewer comments into a roadmap"        → revision-coach mode
"Write an abstract for this paper"                    → abstract-only mode
"Turn this into a literature review paper"            → lit-review mode
"Convert to LaTeX" / "Convert citations to IEEE"      → format-convert mode
"Check citations"                                     → citation-check mode
"Generate an AI disclosure statement for NeurIPS"     → disclosure mode
"Audit my rebuttal draft against the reviews"         → rebuttal-audit mode
Academic Paper Reviewer (6 modes)
"Review this paper"                                   → full mode (Journal-Fit Reviewer + R1/R2/R3 + Devil's Advocate)
"Quick assessment of this paper"                      → quick mode
"Guide me to improve this paper"                      → guided mode
"Check the methodology"                               → methodology-focus mode
"Verify the revisions"                                → re-review mode
"Calibrate this reviewer against my gold set"         → calibration mode
Academic Pipeline (Orchestrator)
"I want to write a complete research paper"           → full pipeline from Stage 1
"I already have a paper, review it"                   → mid-entry at Stage 2.5 (integrity first)
"I received reviewer comments"                        → mid-entry at Stage 4

Pipeline ends with Stage 6: Process Summary — auto-generates a paper creation process record with 6-dimension Collaboration Quality Evaluation (1–100 scoring).

Supported Languages

  • Traditional Chinese (繁體中文) — default when user writes in Chinese
  • English — default when user writes in English
  • Bilingual abstracts (Chinese + English) for academic papers

Using a different language? Socratic mode (deep-research) and Plan mode (academic-paper) use intent-based activation — they detect the meaning of your request, not specific keywords. This means they work in any language without modification.

However, the general Trigger Keywords section (which determines whether the skill is activated at all) still lists English and Traditional Chinese keywords. If you find the skill isn't activating reliably in your language, you can add your language's keywords to the ### Trigger Keywords section in each SKILL.md file to improve matching confidence.

Supported Citation Formats

  • APA 7.0 (default, including Chinese citation rules)
  • Chicago (Notes & Author-Date)
  • MLA
  • IEEE
  • Vancouver

Supported Paper Structures

  • IMRaD (empirical research)
  • Thematic Literature Review
  • Theoretical Analysis
  • Case Study
  • Policy Brief
  • Conference Paper

Skill Details

Per-agent responsibilities and per-stage artifacts now live in docs/ARCHITECTURE.md. Version numbers are anchored here so release metadata stays in one place.

Deep Research (v2.12.1)

13-agent research team. Modes: full, quick, review, lit-review, three-way-scan, fact-check, socratic, systematic-review. Full agent roster and artifacts: see ARCHITECTURE.md §3.

Academic Paper (v3.3.1)

12-agent paper writing pipeline. Modes: full, plan, outline-only, revision, revision-coach, abstract-only, lit-review, format-convert, citation-check, disclosure, rebuttal-audit. Output: MD + DOCX (via Pandoc when available) + LaTeX (APA 7.0 apa7 class / IEEE / Chicago) → PDF via tectonic. Full agent roster and per-phase responsibilities: see ARCHITECTURE.md §3.

Academic Paper Reviewer (v1.11.1)

7-agent multi-perspective review with criterion-bound narrative judgements. Modes: full, re-review, quick, methodology-focus, guided, calibration. Current live reviews and Schema 6 packages remain NOT_CALIBRATED; full calibration can produce a bounded candidate profile, but application to a live review is not wired. No numerical total is mapped to Accept, Minor Revision, Major Revision, or Reject. First-round review panel vs. contract-governed re-review dispatch boundary: see ARCHITECTURE.md §3 Stage 3 / Stage 3'.

Academic Pipeline (v3.22.2)

10-stage orchestrator with integrity verification, two-stage review, Socratic coaching, and collaboration evaluation. Pipeline rules (protocol the agents follow, not runtime guarantees): every stage requires user confirmation checkpoint; integrity verification (Stage 2.5 + 4.5) is MANDATORY with no unrecorded bypass (every override requires user reasoning recorded for Stage 6); R&R Traceability Matrix (Schema 11) maps each reviewer concern to the author's revision claim and records whether the re-review verified it. v3.4 added the Compliance Agent (PRISMA-trAIce + RAISE) at Stage 2.5 / 4.5. v3.5 adds the Collaboration Depth Observer (collaboration_depth_agent, advisory only — never blocks) at every FULL/SLIM checkpoint and at pipeline completion. MANDATORY integrity gates (2.5 / 4.5) explicitly skip the observer so compliance checks are not diluted. Based on Wang & Zhang (2026), IJETHE 23:11. Stage-by-stage matrix with agents, artifacts, and gates: see ARCHITECTURE.md §3.


v3.0 Optimizations: What We Discovered About AI's Structural Limits

What happened

While using ARS to write a reflection article about AI in higher education, I ran into three structural problems that no amount of prompt engineering could fix:

  1. Frame-lock: I asked the AI to run a devil's advocate debate against its own thesis. It did — four rounds, each more refined than the last. But every round stayed inside the frame I'd set. The DA attacked arguments, never premises. It never asked "are we even discussing the right question?" This is the same pattern that caused the 31% citation error rate in v2.7's stress test: the verifying AI and the generating AI share the same cognitive frame.

  2. Sycophancy under pushback: Every time I challenged the DA's attacks, it conceded too quickly. It retracted findings faster than it launched them. The model's training rewards conversational harmony — so "the user pushed back" was treated as evidence that the attack was wrong, when often it just meant the user was persistent.

  3. Intent misdetection: The Socratic Mentor kept trying to converge and produce deliverables ("Want me to write this up?") when I was still exploring. It couldn't distinguish "the user wants a deep philosophical discussion" from "the user wants an RQ brief." Both look like engagement, but they need opposite AI behaviors.

What we changed (v3.0)

Devil's Advocate — Concession Threshold Protocol (deep-research + academic-paper-reviewer)

  • DA must now score every rebuttal on a 1-5 scale before responding
  • Concession only allowed at score ≥4 (rebuttal directly addresses core attack with evidence)
  • Score ≤3: hold position and restate the original attack
  • Anti-sycophancy rules: no consecutive concessions, concession rate tracking, frame-lock detection after each checkpoint

Socratic Mentor — Intent Detection Layer (deep-research)

  • Classifies user intent as exploratory vs. goal-oriented at dialogue start and every 3 turns
  • Exploratory mode: disables auto-convergence, raises max rounds to 60, prohibits "want me to summarize?" prompts
  • Goal-oriented mode: standard convergence behavior
  • Anti-premature-closure rules: in exploratory mode, the user decides when to stop

Socratic Mentor — Dialogue Health Indicator (deep-research)

  • Silent self-assessment every 5 turns on three dimensions: persistent agreement, conflict avoidance, premature convergence
  • Auto-injects challenging questions when agreement pattern detected
  • Invisible to user (to prevent gaming), but log available for post-session review

Why this matters

These optimizations don't solve AI's structural limits — they make the limits visible and manageable. The DA will still eventually concede if pushed hard enough. The Socratic Mentor will still have some convergence bias. But now there are explicit checkpoints that slow down the sycophancy, force the DA to justify concessions, and prevent the Mentor from wrapping up before the user is ready.

The deeper lesson: AI literacy isn't about learning to use AI as a tool, following ethics rules, or fearing AI risks. It's about engaging AI deeply enough to discover its structural limits yourself — and your own thinking limits in the process.


License

This work is licensed under CC-BY-NC 4.0.

You are free to:

  • Share — copy and redistribute the material
  • Adapt — remix, transform, and build upon the material

Under the following terms:

  • Attribution — You must give appropriate credit
  • NonCommercial — You may not use the material for commercial purposes

Attribution format:

Based on Academic Research Skills by Cheng-I Wu
https://github.com/Imbad0202/academic-research-skills

Contributors

Cheng-I Wu (吳政宜) — Author and maintainer

aspi6246 — Contributor. The v3.1 optimization was inspired by patterns from Claude-Code-Skills-for-Academics: read-only constraint pattern, anti-pattern codification as first-class design, cognitive framework approach (teaching "how to think" not just procedures), and lean skill size philosophy.

mchesbro1 — Contributor. Originally proposed and drafted the IS Basket of 8 journals for academic-paper-reviewer/references/top_journals_by_field.md (Issue #5).

cloudenochcsis — Contributor. Extended the IS section from the Basket of 8 to the full Senior Scholars' Basket of 11 — adding Decision Support Systems, Information & Management, and Information and Organization (Issue #7, PR #8). Sourced from the AIS Senior Scholars' List of Premier Journals.

eltociear (Ikko Eltociear Ashimine) — Contributor. Translated the Japanese README (README.ja-JP.md) (PR #161).

xpfo-go (xpfo) — Contributor. Translated the Simplified Chinese README (README.zh-CN.md) (PR #181).

devCharlotte — Contributor. Translated the Korean README (README.ko-KR.md) (PR #469).

Yaobin29 — Contributor. Proposed reviewer-response tooling in PR #433; the deep-research three-way-scan mode and the academic-paper rebuttal-audit mode (rescued from the PR's audit concept) were integrated from that contribution in v3.12.1.

ktao732084-arch — Contributor. Expanded the academic-paper disclosure system with nine medical-publishing policy targets, target-specific required-fact intake, and fail-closed standalone rendering (Issue #596, PR #599); expanded the EQUATOR clinical-reporting reference with condensed CARE, STARD and TRIPOD+AI guidance plus a fail-closed study-design routing sequence (Issue #594, PR #601); and designed and contributed the standalone Chinese-literature resolver, API protocol, and synthetic transport-fixture suite (Issue #595, PR #600).

didacrios — Contributor. Translated the Spanish README (README.es-ES.md).


Changelog

Only the three most recent releases are summarized here. The full release history, including the one-paragraph summaries that used to live on this page, is in CHANGELOG.md.

v3.22.2 (2026-09-25) — Run ledger and handoff check, acronym check, a wider instruction/data boundary, and routing and front-page repairs

Two deterministic checks pinned by synthetic tests; the prompt-level changes are unmeasured: v3.22.2 adds a run ledger (#887). When a pipeline run has a passport file, the orchestrator appends the user's initial instructions, each checkpoint's question and the user's answer in their exact words, step receipts, and file hashes to a local ledger beside the passport. After compaction, resume, or a subagent return, scripts/run_ledger.py report compares the ledger with what the summary or report claims and lists the differences; it now prints that handoff check itself, in English or Traditional Chinese, and hashes the files an entry names when the entry is written (#898). The ledger holds the user's exact words, and docs/DATA_FLOWS.md lists it and how to delete it. The release also adds scripts/check_acronyms.py (#849, proposed by @reiropke), which reports acronyms that are undefined, defined after first use, or defined twice, without calling a model; the prompts have the calling session run it on saved drafts and abstracts, and a review appends its report to the Editorial Decision Letter as an advisory attachment that the decision, the revision roadmap, and re-review criteria do not draw on. Synthetic tests pin both scripts; whether runs write the ledger entries or call the check is unmeasured. The instruction/data boundary now reaches third-party text in dispatches and passport imports (#890), text a receiver reads through its own tool calls, and each skill's main session (#894); a lint pins every copy, and the effect is unmeasured (the opt-in claim-audit judge prompt changes with it, so cached verdicts from the old prompt are not reused). Repairs: the routing core reaches plugin and skills-copy installs (#892); an explicit request stays explicit when the mode's usual input is missing (#889); /ars-lit-review no longer points a run under way to another workflow (#897); the revision coach keeps peer review out of the committee-correspondence variant (#854); the orchestrator scopes "authoritative" skill output to deliverable ownership (#888); and the front page and showcase now match their sources (#908). The routing results come from one session per fixture, a smoke test rather than a rate. One new schema describes the ledger; no existing schema, command model, or effort setting changes.

v3.22.1 (2026-09-23) — Model currency for Claude Opus 5.5, citation-check loading and Chinese APA 7 repairs, and a Pi wrapper fix

Model currency and repairs; the new prompt-level guard is unmeasured: v3.22.1 names Claude Opus 5.5 beside Claude Fable 5.1 as a supported session model, after a two-reader audit of the Opus 5.5 system card that retires no guardrail (#883). The docs add effort guidance (Claude Code starts Opus 5.5 at medium; heavy runs should use high or above), one list-price re-derivation for both models, and tiering guidance that reads the ladder as the vendor's lineup order, not a capability order. Because the card reports that Opus 5.5 acts more often than earlier models on instructions inside pasted text, the revision coach now treats pasted reviewer and committee text as data, pinned by a lint; the effect of that prompt-level guard is unmeasured. The release also repairs mode loading and citation checks: the 13 plugin mode commands invoke their namespaced core skill and resolve bundled references from the plugin root, which restores citation-check loading (#857); Chinese APA 7 checks catch a missing in-text author abbreviation, keep ambiguity exceptions and complete reference-list author fields, and propose a reorder only with evidence of a stroke-order inversion (#882); citation checks now distinguish visible syntax errors from unverified resolution or source claims (#882); and new English, Traditional Chinese, and Korean trigger phrases route citation-check requests, with CI bounding each skill description at 1,024 code points (#858, #864). The Pi wrapper accepts string-array system prompts (#880). No schema, command model, or effort setting changes.

v3.22.0 (2026-09-16) — Output-language-pair contract, locale track, plugin eval suites, and Windows / transport repairs

Additive structure, bounded evidence: v3.22.0 lets a run declare its output language pair through a registry-keyed Schema 4 field whose absence reproduces the legacy files exactly (#862 Phase 1, PR #869), and stands up the locale track around it: an es-ES README and conservative trigger phrases contributed by @didacrios, and a community-maintained locale-pack policy with a provisional single-owner route. Two claude plugin eval suites (revision-coach, citation-check) and the reviewer-calibration harness ship as regression guards and dispatch substrates only; none claims a measured uplift or calibration value. Repairs: /ars-mark-read and the other five lock sites run on Windows through one shared helper with an msvcrt backend, the OpenAI request builders stop sending parameters GPT-6 Astra rejects, the contained Codex transport refuses effort=ultra, audit provenance records the actual judge identity, Socratic path F6 no longer preselects a direction, and an unsupported claim can no longer be rescued by hedging. READMEs keep three releases; Gartenberg et al. and Wang, Li et al. join the human-in-the-loop anchors. Roadmap Phase 4 (stage-level evidence ceilings) is not delivered in this release; its window carries forward.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

45 total
  1. > **Two deterministic checks pinned by synthetic tests; the prompt-level changes are unmeasured:** v3.22.2 adds a run ledger (#887). When a pipeline run has a passport file, the orchestrator appends the user's initial instructions, each checkpoint's question and the user's answer in their exact words, step receipts, and file hashes to a local ledger beside the passport. After compaction, resume, or a subagent return, `scripts/run_ledger.py report` compares the ledger with what the summary or report claims and lists the differences; it now prints that handoff check itself, in English or Traditional Chinese, and hashes the files an entry names when the entry is written (#898). The ledger holds the user's exact words, and `docs/DATA_FLOWS.md` lists it and how to delete it. The release also adds `scripts/check_acronyms.py` (#849, proposed by @reiropke), which reports acronyms that are undefined, defined after first use, or defined twice, without calling a model; the prompts have the calling session run it on saved drafts and abstracts, and a review appends its report to the Editorial Decision Letter as an advisory attachment that the decision, the revision roadmap, and re-review crite

  2. > **Model currency and repairs; the new prompt-level guard is unmeasured:** v3.22.1 names Claude Opus 5.5 beside Claude Fable 5.1 as a supported session model, after a two-reader audit of the Opus 5.5 system card that retires no guardrail (#883). The docs add effort guidance (Claude Code starts Opus 5.5 at `medium`; heavy runs should use `high` or above), one list-price re-derivation for both models, and tiering guidance that reads the ladder as the vendor's lineup order, not a capability order. Because the card reports that Opus 5.5 acts more often than earlier models on instructions inside pasted text, the revision coach now treats pasted reviewer and committee text as data, pinned by a lint; the effect of that prompt-level guard is unmeasured. The release also repairs mode loading and citation checks: the 13 plugin mode commands invoke their namespaced core skill and resolve bundled references from the plugin root, which restores citation-check loading (#857); Chinese APA 7 checks catch a missing in-text author abbreviation, keep ambiguity exceptions and complete reference-list author fields, and propose a reorder only with evidence of a stroke-order inversion (#882); citation

  3. > **Additive structure, bounded evidence:** v3.22.0 lets a run declare its output language pair through a registry-keyed Schema 4 field whose absence reproduces the legacy files exactly (#862 Phase 1, PR #869), and stands up the locale track around it: an es-ES README and conservative trigger phrases contributed by @didacrios, and a community-maintained locale-pack policy with a provisional single-owner route. Two `claude plugin eval` suites (revision-coach, citation-check) and the reviewer-calibration harness ship as regression guards and dispatch substrates only; none claims a measured uplift or calibration value. Repairs: `/ars-mark-read` and the other five lock sites run on Windows through one shared helper with an `msvcrt` backend, the OpenAI request builders stop sending parameters GPT-6 Astra rejects, the contained Codex transport refuses `effort=ultra`, audit provenance records the actual judge identity, Socratic path F6 no longer preselects a direction, and an unsupported claim can no longer be rescued by hedging. READMEs keep three releases; Gartenberg et al. and Wang, Li et al. join the human-in-the-loop anchors. Roadmap Phase 4 (stage-level evidence ceilings) is not de

  4. ## What's Changed - Apply Chinese-aware title matching in the four index resolvers — #799 - Declare the markdown-it-py floor and make the autolink round-trip tail run visibly — #802 - Fail `check_surface_form_parity` with an environment error when pyyaml is missing — #803 - Strip CJK outer wrapper marks only when they enclose one balanced unit — #804 - Fix a duplicated word in the MLA citation key rules — #805 (by @LeslieLi46) - Add the skill-inventory parity lint — #810 - September harness-retirement audit, zero findings — #812 - De-stale the RISK_REGISTER R10 residual gap — #814 - Align docs and guardrails to Claude Fable 5.1 and GPT-6 Astra after a full read of both system cards — #820 - Release preparation and version alignment — #822 ## Highlights - Names Claude Fable 5.1 as the current frontier model and lists `gpt-6-astra` as a provisional cross-model verifier on both transports; the OpenAI recommendation moves to the current generation under the existing generation-currency policy, while `gpt-5.6-sol` keeps its transport-qualified validated status. The contained Codex citation transport's closed reasoning-effort set gains `ultra`. - Adds checkpoint decision provenance: o

  5. ## What's Changed - Add OrcaRouter to the community directory — #782 - Update cross-model recommendation surfaces for generation currency — #784 - Repair the Codex ChatGPT-subscription citation transport for codex-cli 0.147.0 — #786 - Record the first Promotion Bakeoff and validate `gpt-5.6-sol` for that transport — #788 - Consolidate Markdown parsing helpers and harden link/heading semantics — #791 - Register write-scope guard launcher degradations — #792 - Re-derive `data_access_level` for standalone paper and reviewer skills — #793 - Add sealed bakeoff and default-off workflow-profile contracts — #795 - Add the opt-in inquiry ledger, alternative-register design freeze, and bounded review-criteria proving set — #796 - Release preparation and version alignment — #797 ## Highlights - Repairs the contained ChatGPT-subscription citation transport for Codex CLI 0.147.0, covering stderr authentication attestation, provider-incompatible schema keywords, and the code-mode-host/web-search interaction. - Adds sealed preregistration for future Promotion Bakeoffs and records the first counterbalanced 180-call bakeoff validating `gpt-5.6-sol` for the ChatGPT-subscription citation transport

Code frequency

additions and deletions
+152.4K-152.4KWeek of 2026-02-22: +22,985 linesWeek of 2026-02-22: -352 linesWeek of 2026-03-01: +4,221 linesWeek of 2026-03-01: -574 linesWeek of 2026-03-08: +18,137 linesWeek of 2026-03-08: -10,393 linesWeek of 2026-03-15: +0 linesWeek of 2026-03-15: -0 linesWeek of 2026-03-22: +960 linesWeek of 2026-03-22: -269 linesWeek of 2026-03-29: +908 linesWeek of 2026-03-29: -41 linesWeek of 2026-04-05: +6,003 linesWeek of 2026-04-05: -3,786 linesWeek of 2026-04-12: +3,748 linesWeek of 2026-04-12: -1,419 linesWeek of 2026-04-19: +18,433 linesWeek of 2026-04-19: -1,025 linesWeek of 2026-04-26: +17,950 linesWeek of 2026-04-26: -472 linesWeek of 2026-05-03: +17,591 linesWeek of 2026-05-03: -303 linesWeek of 2026-05-10: +24,584 linesWeek of 2026-05-10: -459 linesWeek of 2026-05-17: +18,586 linesWeek of 2026-05-17: -767 linesWeek of 2026-05-24: +9,968 linesWeek of 2026-05-24: -239 linesWeek of 2026-05-31: +18,786 linesWeek of 2026-05-31: -1,357 linesWeek of 2026-06-07: +22,873 linesWeek of 2026-06-07: -1,097 linesWeek of 2026-06-14: +4,453 linesWeek of 2026-06-14: -309 linesWeek of 2026-06-21: +1,063 linesWeek of 2026-06-21: -6 linesWeek of 2026-06-28: +5,448 linesWeek of 2026-06-28: -524 linesWeek of 2026-07-05: +1,912 linesWeek of 2026-07-05: -56 linesWeek of 2026-07-12: +12,643 linesWeek of 2026-07-12: -490 linesWeek of 2026-07-19: +44,706 linesWeek of 2026-07-19: -375 linesWeek of 2026-07-26: +39,186 linesWeek of 2026-07-26: -2,639 linesWeek of 2026-08-02: +71,555 linesWeek of 2026-08-02: -787 linesWeek of 2026-08-09: +152,354 linesWeek of 2026-08-09: -5,653 linesWeek of 2026-08-16: +13,605 linesWeek of 2026-08-16: -566 linesWeek of 2026-08-23: +13,739 linesWeek of 2026-08-23: -320 linesWeek of 2026-08-30: +935 linesWeek of 2026-08-30: -21 linesWeek of 2026-09-06: +21,831 linesWeek of 2026-09-06: -441 linesFeb 22, 2026Sep 6, 2026
+589.2K lines added, -34.7K removed over the last year.

Commits per week

last 52 weeks
1530Week of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 6 commitsWeek of 2026-03-01: 9 commitsWeek of 2026-03-08: 33 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 13 commitsWeek of 2026-03-29: 7 commitsWeek of 2026-04-05: 33 commitsWeek of 2026-04-12: 53 commitsWeek of 2026-04-19: 153 commitsWeek of 2026-04-26: 16 commitsWeek of 2026-05-03: 21 commitsWeek of 2026-05-10: 23 commitsWeek of 2026-05-17: 33 commitsWeek of 2026-05-24: 27 commitsWeek of 2026-05-31: 43 commitsWeek of 2026-06-07: 48 commitsWeek of 2026-06-14: 13 commitsWeek of 2026-06-21: 3 commitsWeek of 2026-06-28: 12 commitsWeek of 2026-07-05: 8 commitsWeek of 2026-07-12: 25 commitsWeek of 2026-07-19: 14 commitsWeek of 2026-07-26: 19 commitsWeek of 2026-08-02: 28 commitsWeek of 2026-08-09: 44 commitsWeek of 2026-08-16: 22 commitsWeek of 2026-08-23: 7 commitsWeek of 2026-08-30: 4 commitsWeek of 2026-09-06: 12 commitsWeek of 2026-09-13: 16 commitsWeek of 2026-09-20: 17 commitsSep 27, 2025Sep 20, 2026
762 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 12 commitsSun 1:00 — 5 commitsSun 2:00 — 1 commitsSun 3:00 — 2 commitsSun 4:00 — 1 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 1 commitsSun 8:00 — 5 commitsSun 9:00 — 0 commitsSun 10:00 — 8 commitsSun 11:00 — 6 commitsSun 12:00 — 4 commitsSun 13:00 — 6 commitsSun 14:00 — 6 commitsSun 15:00 — 3 commitsSun 16:00 — 5 commitsSun 17:00 — 2 commitsSun 18:00 — 5 commitsSun 19:00 — 0 commitsSun 20:00 — 3 commitsSun 21:00 — 0 commitsSun 22:00 — 3 commitsSun 23:00 — 3 commitsMon 0:00 — 6 commitsMon 1:00 — 1 commitsMon 2:00 — 1 commitsMon 3:00 — 2 commitsMon 4:00 — 1 commitsMon 5:00 — 2 commitsMon 6:00 — 4 commitsMon 7:00 — 1 commitsMon 8:00 — 0 commitsMon 9:00 — 7 commitsMon 10:00 — 10 commitsMon 11:00 — 14 commitsMon 12:00 — 10 commitsMon 13:00 — 9 commitsMon 14:00 — 7 commitsMon 15:00 — 10 commitsMon 16:00 — 6 commitsMon 17:00 — 12 commitsMon 18:00 — 4 commitsMon 19:00 — 8 commitsMon 20:00 — 1 commitsMon 21:00 — 4 commitsMon 22:00 — 3 commitsMon 23:00 — 2 commitsTue 0:00 — 2 commitsTue 1:00 — 0 commitsTue 2:00 — 2 commitsTue 3:00 — 1 commitsTue 4:00 — 1 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 3 commitsTue 9:00 — 5 commitsTue 10:00 — 9 commitsTue 11:00 — 4 commitsTue 12:00 — 4 commitsTue 13:00 — 1 commitsTue 14:00 — 3 commitsTue 15:00 — 6 commitsTue 16:00 — 3 commitsTue 17:00 — 11 commitsTue 18:00 — 7 commitsTue 19:00 — 1 commitsTue 20:00 — 3 commitsTue 21:00 — 2 commitsTue 22:00 — 2 commitsTue 23:00 — 3 commitsWed 0:00 — 1 commitsWed 1:00 — 1 commitsWed 2:00 — 1 commitsWed 3:00 — 3 commitsWed 4:00 — 2 commitsWed 5:00 — 2 commitsWed 6:00 — 1 commitsWed 7:00 — 0 commitsWed 8:00 — 4 commitsWed 9:00 — 4 commitsWed 10:00 — 7 commitsWed 11:00 — 14 commitsWed 12:00 — 7 commitsWed 13:00 — 15 commitsWed 14:00 — 7 commitsWed 15:00 — 5 commitsWed 16:00 — 22 commitsWed 17:00 — 10 commitsWed 18:00 — 4 commitsWed 19:00 — 7 commitsWed 20:00 — 7 commitsWed 21:00 — 2 commitsWed 22:00 — 5 commitsWed 23:00 — 2 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 2 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 1 commitsThu 7:00 — 2 commitsThu 8:00 — 1 commitsThu 9:00 — 6 commitsThu 10:00 — 14 commitsThu 11:00 — 26 commitsThu 12:00 — 17 commitsThu 13:00 — 15 commitsThu 14:00 — 19 commitsThu 15:00 — 16 commitsThu 16:00 — 21 commitsThu 17:00 — 18 commitsThu 18:00 — 19 commitsThu 19:00 — 4 commitsThu 20:00 — 2 commitsThu 21:00 — 1 commitsThu 22:00 — 8 commitsThu 23:00 — 5 commitsFri 0:00 — 3 commitsFri 1:00 — 0 commitsFri 2:00 — 1 commitsFri 3:00 — 2 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 1 commitsFri 7:00 — 1 commitsFri 8:00 — 1 commitsFri 9:00 — 5 commitsFri 10:00 — 1 commitsFri 11:00 — 5 commitsFri 12:00 — 4 commitsFri 13:00 — 4 commitsFri 14:00 — 5 commitsFri 15:00 — 1 commitsFri 16:00 — 1 commitsFri 17:00 — 5 commitsFri 18:00 — 8 commitsFri 19:00 — 0 commitsFri 20:00 — 2 commitsFri 21:00 — 2 commitsFri 22:00 — 3 commitsFri 23:00 — 4 commitsSat 0:00 — 0 commitsSat 1:00 — 2 commitsSat 2:00 — 0 commitsSat 3:00 — 1 commitsSat 4:00 — 1 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 1 commitsSat 8:00 — 1 commitsSat 9:00 — 5 commitsSat 10:00 — 4 commitsSat 11:00 — 12 commitsSat 12:00 — 5 commitsSat 13:00 — 22 commitsSat 14:00 — 11 commitsSat 15:00 — 4 commitsSat 16:00 — 3 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 3 commitsSat 20:00 — 4 commitsSat 21:00 — 1 commitsSat 22:00 — 5 commitsSat 23:00 — 8 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits760 (96%)
Community commits35 (4%)

795 commits in total over the last year.

DateListRankStars gained
Sep 10, 2026weekly#11+2,430
Sep 9, 2026weekly#11+2,430
Sep 8, 2026weekly#4+2,406
Sep 7, 2026weekly#13+2,334
Sep 4, 2026daily#11+799
Sep 3, 2026daily#11+799
Sep 2, 2026daily#2+193
Sep 1, 2026daily#2+161
Jul 12, 2026daily#22+2
Jul 8, 2026daily#25+4
Jun 9, 2026daily#22+18
Jun 1, 2026daily#22+31
May 31, 2026daily#25+57
May 30, 2026daily#14+80
May 26, 2026daily#20+50