yaojingang/yao-meta-skillPublic

YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

AI summary: A comprehensive governance framework for creating, evaluating, and releasing reusable AI agent skills.

Stars
2.7K
+1 today
Forks
253
Watchers
10
Open issues
2
Open PRs
1
Contributors
~3
Commits
151
Branches
9

PythonMITCreated Mar 31, 2026Last push 1mo ago+58 stars this week+103 this month

Quick answers

What is yao-meta-skill?
A comprehensive governance framework for creating, evaluating, and releasing reusable AI agent skills.
What does yao-meta-skill do?
Yao Meta Skill provides a complete, highly rigorous lifecycle management system for transforming rough prompts and workflows into robust, highly reusable AI agent skills. It entirely replaces manual, conversational skill creation with a strict engineering pipeline involving explicit inputs, target compilers, and strict quality evaluation gates. The framework heavily enforces an evidence-based release process, strictly requiring execution proof, blind-review packs, and native permission probes before any skill can be published. Furthermore, it incorporates an advanced operating loop for post-release maintenance, utilizing metadata-only telemetry and precise adoption drift reports to intelligently guide future updates.
Who is yao-meta-skill for?
This framework is designed for AI tool creators, prompt engineers, and engineering managers who need a highly structured, auditable way to build complex AI agent skills. It is essential for teams prioritizing strict governance, cross-platform portability, and rigorous evaluation.
How do I get started with yao-meta-skill?
Please refer to the repository README to begin utilizing the Skill OS commands.
How popular is yao-meta-skill on GitHub?
yaojingang/yao-meta-skill has 2,688 stars and 253 forks on GitHub, and gained 58 stars in the last 7 days.
What license does yao-meta-skill use?
yaojingang/yao-meta-skill is released under the MIT license.

Star history

since Jul 29, 2026
01K2KJul 2026Aug 2026Sep 2026Oct 2026
2.7K stars as of Oct 3, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
OctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 17 commits2026-04-01: 13 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 5 commits2026-04-07: 0 commits2026-04-08: 6 commits2026-04-09: 4 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 4 commits2026-04-15: 1 commit2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 1 commit2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 6 commits2026-04-24: 0 commits2026-04-25: 1 commit2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 1 commit2026-05-01: 1 commit2026-05-02: 0 commits2026-05-03: 1 commit2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 1 commit2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 1 commit2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 1 commit2026-06-13: 4 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 1 commit2026-06-18: 3 commits2026-06-19: 1 commit2026-06-20: 2 commits2026-06-21: 1 commit2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 1 commit2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 2 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 1 commit2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 6 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 0 commits2026-08-03: 0 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits2026-08-09: 0 commits2026-08-10: 0 commits2026-08-11: 0 commits2026-08-12: 13 commits2026-08-13: 0 commits2026-08-14: 8 commits2026-08-15: 0 commits2026-08-16: 20 commits2026-08-17: 11 commits2026-08-18: 0 commits2026-08-19: 0 commits2026-08-20: 0 commits2026-08-21: 0 commits2026-08-22: 0 commits2026-08-23: 0 commits2026-08-24: 0 commits2026-08-25: 0 commits2026-08-26: 0 commits2026-08-27: 0 commits2026-08-28: 0 commits2026-08-29: 0 commits2026-08-30: 0 commits2026-08-31: 0 commits2026-09-01: 0 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits2026-09-06: 0 commits2026-09-07: 0 commits2026-09-08: 0 commits2026-09-09: 0 commits2026-09-10: 0 commits2026-09-11: 0 commits2026-09-12: 0 commits2026-09-13: 0 commits2026-09-14: 0 commits2026-09-15: 0 commits2026-09-16: 0 commits2026-09-17: 0 commits2026-09-18: 0 commits2026-09-19: 0 commits2026-09-20: 0 commits2026-09-21: 0 commits2026-09-22: 0 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits2026-09-27: 0 commits2026-09-28: 0 commits2026-09-29: 0 commits2026-09-30: 0 commits2026-10-01: 0 commits2026-10-02: 0 commits2026-10-03: 0 commits
138 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    MIT

What yao-meta-skill does

Yao Meta Skill provides a complete, highly rigorous lifecycle management system for transforming rough prompts and workflows into robust, highly reusable AI agent skills. It entirely replaces manual, conversational skill creation with a strict engineering pipeline involving explicit inputs, target compilers, and strict quality evaluation gates. The framework heavily enforces an evidence-based release process, strictly requiring execution proof, blind-review packs, and native permission probes before any skill can be published. Furthermore, it incorporates an advanced operating loop for post-release maintenance, utilizing metadata-only telemetry and precise adoption drift reports to intelligently guide future updates.

This framework is designed for AI tool creators, prompt engineers, and engineering managers who need a highly structured, auditable way to build complex AI agent skills. It is essential for teams prioritizing strict governance, cross-platform portability, and rigorous evaluation.

  • Skill IR compilation architecture: Uses an advanced intermediate representation and target compilers to deliver skills consistently across vastly different agent platforms.
  • Evidence-based release gates: Strictly blocks publication until execution evidence, blind-review packs, and complex reproducibility manifests are fully validated.
  • Inference-first intent dialogue: Captures deep skill requirements through a highly structured, two-round dialogue that strictly records assumptions.
  • Runtime permission probing: Deeply analyzes packaged target adapters for explicit permission requirements and strict native enforcement flags.
  • Metadata-only telemetry tracking: Tracks skill adoption and severe drift using strictly privacy-preserving client events without storing any raw content.

Where teams use it

Standardizing team agent skills

Engineering teams actively upgrade personal AI workflows into fully governed assets with highly explicit interface contracts and automated release notes.

Ensuring public release quality

Maintainers strictly verify that a newly developed agent skill meets world-class evidence standards before making sweeping public claims.

Cross-platform skill deployment

Developers efficiently write a complex skill once and successfully compile it for multiple diverse agent targets, including Claude, OpenAI, and VS Code.

Monitoring skill adoption drift

Operators leverage advanced metadata telemetry to quickly detect when a skill's real-world usage significantly diverges from its original intent.

Getting started: Please refer to the repository README to begin utilizing the Skill OS commands.

README

main branch

Yao Meta Skill

CI License: MIT English 中文 日本語 Français Русский

YAO stands for Yielding AI Outcomes: the goal is not to generate more prompt text, but to produce reusable AI assets and real operational outcomes.

yao-meta-skill creates, evaluates, packages, and governs reusable agent skills. The 1.0 line focused on turning repeated workflows into installable, readable, cross-platform skill packages. The 2.0 line expands that factory into a Skill OS: a governed system for modeling a skill once, compiling it for multiple targets, testing its behavior, reviewing its release evidence, and tracking the next iteration.

Quick Start · Skill OS 2.0 · 1.0 vs 2.0 · Operator UX · Benchmark · Examples · Evals · Failure Library · Method Doctrine

Skill OS 2.0 Upgrade

Skill OS 2.0 keeps the original promise of yao-meta-skill, but makes the package lifecycle more explicit. Instead of stopping at SKILL.md, it adds a semantic contract, target compilers, evaluation evidence, release gates, and operation reports around the skill.

  • Skill IR: a platform-neutral intermediate representation for intent, triggers, inputs, outputs, boundaries, references, and expected artifacts.
  • Target compilers and adapters: generated surfaces for OpenAI, Claude, generic agent skills, Agent Skills compatible packages, and VS Code-oriented workflows.
  • Output Eval Lab: trigger checks, output assertions, execution evidence, timing and token evidence, benchmark reproducibility, blind-review packs, answer keys, and adjudication reports.
  • Review Studio 2.0: a single HTML gate page for intent, triggers, output eval, context cost, runtime checks, trust, Skill Atlas signals, adoption drift, waivers, annotations, release evidence, warnings, blockers, and fix actions.
  • Evidence and release governance: evidence consistency checks, package verification, install simulation, runtime permission probes, world-class evidence intake, world-class ledger, operator runbook, and public claim guard.
  • SkillOps loop: metadata-only adoption drift, telemetry hooks, adaptive proposals, daily and weekly curator reports, and portfolio-level drift signals.

Current posture: the repository is ready for beta and external testing, while stronger public "world-class" claims remain evidence-gated. Provider-backed production evidence, human blind-review evidence, native permission execution, and real-client telemetry are tracked as separate evidence tasks instead of being treated as completed work.

See the companion artifacts:

From 1.0 to 2.0

Dimension 1.0 focus 2.0 upgrade
Product role Create, refactor, evaluate, and package reusable skills. Govern the full lifecycle of a skill: creation, compilation, evaluation, review, release, telemetry, and iteration.
Architecture SKILL.md, agents/interface.yaml, manifest files, and report artifacts. Skill IR, target compilers, adapters, gate contracts, evidence ledgers, release locks, and action-oriented review pages.
Cross-platform delivery OpenAI, Claude, and generic package targets. Adds broader Agent Skills and VS Code-oriented compatibility, with registry-readable compatibility records.
Quality model Trigger and structure checks plus report-based review. Output eval, benchmark reproducibility, execution evidence, failure disclosure, blind-review packs, and evidence consistency checks.
Report experience Overview HTML and first-pass review pages. Bilingual Skill Overview v2, Review Studio 2.0, reviewer annotations, action cards, charts, and audit-oriented report contracts.
Release boundary Package output with basic validation. Package verification, install simulation, runtime permission probes, release locks, public claim guard, and operator runbooks.
Operating loop Manual feedback and local iteration. Adoption drift, metadata telemetry, SkillOps reports, adaptive proposals, and portfolio-level drift detection.

2.0 Use Cases

  • Create a new skill from repeated work: start with a workflow note, prompt set, transcript, runbook, or document pattern, then generate a package with a lean entrypoint, explicit inputs and outputs, references, reports, and the lightest justified gates.
  • Upgrade a personal skill into a team asset: add interface contracts, manifests, target adapters, trust checks, output evals, reviewer waivers, release notes, and Review Studio evidence before other people depend on the skill.
  • Prepare a skill for beta release: run package verification, install simulation, compatibility checks, runtime permission probes, and evidence consistency checks, then separate beta readiness from stronger public claims.
  • Keep a skill useful after release: use metadata-only telemetry, adoption drift, feedback logs, SkillOps reports, and adaptive proposals to decide whether the next move should be documentation, an eval, a skill patch, or a governance update.
  • Compare with other meta-skill approaches: keep Anthropic/OpenAI-style conversational creation and lean instruction writing where they fit, then use yao-meta-skill when the package needs evidence, portability, release gates, and repeatable maintenance.

Operator UX Commands

These read-only helper commands turn common maintainer questions into repeatable diagnostics:

python3 scripts/yao.py install-status --expected-source .
python3 scripts/yao.py localized-doc-sync-check
python3 scripts/yao.py pr-review-report 4 --repo yaojingang/yao-meta-skill
  • install-status explains whether the active skill is coming from .codex/skills, .agents/skills, or the disabled mirror, and flags duplicate active installs.
  • localized-doc-sync-check verifies that the Chinese README carries the public homepage sections that were added to the English README.
  • pr-review-report reads GitHub PR metadata, changed files, status checks, and suggested local commands without merging or mutating the PR.

Capability Surface

It turns rough workflows, transcripts, prompts, notes, and runbooks into reusable skill packages with:

  • a clear trigger surface
  • a lean SKILL.md
  • optional references, scripts, and evals
  • an inference-first intent dialogue that asks one personalized question only for core task, output, or direction forks, stops after two rounds, and records structured assumptions without retaining a full transcript
  • a silent-by-default GitHub benchmark scan plus reference synthesis that studies top public repositories and world-class pattern tracks, then surfaces only real conflicts or uncertainty to the user
  • a generated visual HTML overview for each newly initialized skill
  • a Review Studio 2.0 HTML gate page that combines intent, trigger, output eval, context, runtime, trust, atlas, adoption drift, reviewer waivers, reviewer annotations, release evidence, and per-warning fix actions
  • a Skill OS 2.0 audit that maps each world-class requirement to current evidence, human-required gaps, and external-required gaps
  • a Skill OS 2.0 blueprint coverage report that maps the upgrade plan's core modules and recommended PRs to concrete artifacts, commands, and tests
  • a world-class evidence plan that turns remaining provider, human, native-permission, and real-client telemetry gaps into executable evidence tasks
  • a world-class evidence ledger that records which external and human evidence is accepted or still pending without treating planned work as proof
  • a world-class evidence intake contract that validates external and human evidence packets for provenance, privacy, artifact refs, and anti-overclaim rules before ledger review
  • a redacted world-class preflight report that checks local files, environment readiness, human/external prerequisites, and source blockers before operators collect evidence
  • a world-class submission review queue that compares evidence packets, intake validation, source artifacts, and ledger state without accepting evidence
  • a world-class operator runbook that gives reviewers the exact commands, artifacts, and collection checklist needed to close remaining evidence gaps
  • a world-class claim guard that scans public claim surfaces and blocks premature completed/true claims while the evidence ledger still has pending external or human evidence
  • a benchmark reproducibility manifest that checks methodology sections, required artifacts, failure disclosure, and reproduction commands
  • an evidence consistency gate that compares generated reports against each other so benchmark, overview, interpretation, adoption, world-class ledger, coverage, and Review Studio facts do not drift silently
  • Output Eval Lab evidence with assertion grading, execution/timing/token evidence, a blind A/B review pack, a separate answer key, and reviewer adjudication reports
  • a runtime permission probe report that checks packaged target adapters for explicit permission metadata, native-enforcement flags, metadata fallback notes, and residual risks
  • a Python compatibility gate that catches supported-runtime syntax hazards before they reach GitHub Actions or packaged distribution
  • a side-by-side HTML review studio for first-pass human review
  • an artifact design profile that defines visual direction, layout patterns, and quality gates for reports, tutorials, dashboards, screenshots, and review pages
  • a prompt quality profile that abstracts need modeling, RTF mapping, complexity, and quality checks into reviewer-visible evidence instead of bloating SKILL.md
  • a systems-thinking model that maps boundaries, feedback loops, drift risks, recurring failure patterns, and highest-leverage quality moves
  • three high-value next iteration directions after the first package is created
  • a lightweight feedback log that does not require a full promotion cycle
  • a local-first metadata-only adoption and drift report that turns real usage signals into next iteration candidates, with optional yao.py CLI run capture, external client event emit hooks, hook recipes, and JSONL import that record command names and outcomes without arguments or raw content
  • an explicit-source adaptive proposal loop that summarizes redacted repeated user preferences and generates approval-gated adaptation proposals without scanning private logs or writing source files
  • a SkillOps opportunity scorer and decision policy that ranks redacted repeated signals, maps them to report-only, AGENTS update, existing-skill patch, or eval-addition actions, and keeps every durable write approval-gated
  • a weekly SkillOps curator report that aggregates daily opportunities, Skill Atlas portfolio signals, release lock state, and world-class evidence gaps into a proposal-only maintenance queue
  • a Browser/Chrome Native Messaging telemetry host that can receive length-prefixed metadata-only client events and generate a local launcher plus manifest without storing raw content
  • a Skill Atlas drift layer that reads aggregate adoption reports and surfaces portfolio-level drift signals without packaging raw telemetry logs
  • a baseline compare report for with-skill vs baseline review
  • a conversation-style, archetype-aware quickstart that steers new packages toward scaffold, production, library, or governed fits
  • Skill IR as the platform-neutral semantic contract, plus compiler reports and client-specific adapters
  • Registry audit metadata with package version, owner, license, checksum, and compatibility matrix
  • governance, promotion, and portability checks built into the default flow

Architecture

Hero view: Skill OS 2.0 turns messy operational input into a governed, reusable skill package through a model, compile, evaluate, release, and operate loop.

flowchart LR
    A["Inputs<br/>workflow / prompt / transcript / docs / notes"] --> B["Intent model<br/>job / outputs / exclusions / standards"]
    B --> C["Skill IR<br/>trigger / contracts / resources / evidence"]
    C --> D["Skill package<br/>SKILL.md / references / scripts / reports"]
    C --> E["Target compilers<br/>OpenAI / Claude / generic / Agent Skills / VS Code"]
    D --> F["Eval Lab<br/>trigger / output / benchmark / runtime"]
    E --> F
    F --> G["Review Studio<br/>gates / warnings / actions / waivers"]
    G --> H["Release boundary<br/>package verification / install simulation / claim guard"]
    H --> I["SkillOps loop<br/>feedback / adoption drift / next iteration"]
    I --> B
Loading

Read it in 10 seconds:

  • Inputs: start from rough operational material instead of a polished spec.
  • Intent model: make the job, outputs, exclusions, constraints, and standards explicit before generating files.
  • Skill IR: keep the semantic contract separate from any single platform format.
  • Package and compile: generate the lean skill package and the target-specific adapters from the same source model.
  • Evaluate and review: turn trigger behavior, output quality, runtime checks, and trust signals into reviewable evidence.
  • Release and operate: publish only within the current evidence boundary, then feed adoption drift and reviewer feedback into the next iteration.

Weighted Quality Benchmark

This benchmark is a project-level engineering review, scored from 0-10 per dimension and weighted to 100. GitHub stars are intentionally excluded because they measure ecosystem heat, not meta-skill engineering quality.

The score is local engineering evidence, not a claim of world-class readiness. Public superiority claims still depend on accepted external and human evidence in the world-class ledger.

Weighted score formula: sum(score / 10 * weight).

Meta Skill Method Depth 15 Context Discipline 10 Toolchain 15 Eval/Test Rigor 20 Governance 15 Portability 10 Onboarding/Review 5 Local Reliability 10 Weighted Score
Yao Meta Skill 9.5 8.0 9.5 9.5 9.5 9.0 6.5 9.5 91.5
Anthropic Skill Creator 9.0 6.5 8.5 7.5 4.0 5.0 7.5 5.0 67.5
OpenAI Skill Creator 8.5 9.5 5.0 2.0 3.0 4.0 8.5 4.0 50.5
Rank Meta Skill Score Core Positioning
1 Yao Meta Skill 91.5 A complete engineering, evaluation, governance, and portability system for reusable skills.
2 Anthropic Skill Creator 67.5 Strong methodology and iteration loop, with weaker local execution reliability and governance coverage.
3 OpenAI Skill Creator 50.5 Best treated as a concise skill-writing method guide rather than a full engineering system.

Human Blind A/B Review Snapshot

On 2026-06-29, a single human reviewer compared yao-meta-skill with the bundled OpenAI skill-creator across five realistic skill-creation scenarios: support triage, revenue reconciliation, webinar repurposing, incident postmortems, and PR review follow-up. The reviewer confirmed decisions were completed before the answer key was opened.

Result: yao-meta-skill was selected in 5/5 cases.

Evidence:

Boundary: this is single-reviewer blind preference evidence. It is not provider-backed independent model execution evidence, and the per-case rationale fields are still empty.

Best-Fit Scenarios

  • Choose Yao Meta Skill when the target is a reusable team asset with explicit boundaries, trigger evaluation, governance, packaging, portability, and local execution checks.
  • Choose Anthropic Skill Creator when the target is a conversation-first creation loop and the priority is human-guided iteration over repository-level governance.
  • Choose OpenAI Skill Creator when the target is a compact reference for writing lean skill instructions and keeping context small.
  • A practical hybrid pattern is still useful: draft conversationally, then use yao-meta-skill to harden the package, add evidence, and make it team-ready.

Quick Start

Install the skill globally for Codex first:

npx -y skills add yaojingang/yao-meta-skill -a codex -g -y

To install it for every supported agent, replace -a codex with -a '*':

npx -y skills add yaojingang/yao-meta-skill -a '*' -g -y

After installation, restart the client. Then ask for tasks such as "create a skill from this workflow", "improve this existing skill", "evaluate this skill", or "add evals to this skill" to trigger yao-meta-skill.

Update notifications and one-confirmation update

Each activation can run a cached update preflight. It checks the official GitHub VERSION at most once every 24 hours and shows each newer stable version once. Network failures leave the active Skill task unchanged.

python3 scripts/yao.py check-update --notice --self
python3 scripts/yao.py check-update --force --self
python3 scripts/yao.py self-update --self
python3 scripts/yao.py self-update --self --yes

self-update recognizes verified Agent Skills CLI and installed Codex plugin channels. The command displays its plan first and requires --yes before changing an installation. Development checkouts, unmanaged copies, non-official sources, and ambiguous multi-channel installs stay read-only. Restart Codex or the active AI client after a successful update. See Update Delivery for the channel and recovery contract.

  1. Describe the workflow, prompt set, or repeated task you want to turn into a skill.
  2. Start with a short, human intent dialogue so the real job, outputs, exclusions, constraints, and standards are explicit.
  3. Let quickstart clarify intent first, then run silent benchmark scan and reference synthesis; it only surfaces explicit questions when intent is still unclear or when there is a real design conflict.
  4. Use the archetype-aware quickstart or the full authoring flow to generate or improve the package in scaffold, production, library, or governed mode.
  5. Review the generated reports/skill-interpretation.html first for the bilingual interpretation report. It defaults to Simplified Chinese and provides an English switch in the top right. Then open reports/skill-overview.html for the audit scorecard and reports/review-studio.html to inspect release blockers, permission approvals, and evidence paths in one page before adding more structure.

Target-specific CLI commands require an explicit Skill path. Commands that operate on Yao Meta Skill itself also require --self. Runtime update cache and opt-in CLI telemetry use user-level cache/state directories, so external Skill work leaves the Yao source and installed package unchanged.

Or use the unified authoring CLI:

python3 scripts/yao.py quickstart --output-dir .
python3 scripts/yao.py github-benchmark-scan my-skill --query "release workflow portability"
python3 scripts/yao.py reference-scan my-skill \
  --external-reference "World Class Method::method::Borrow a tight evaluation loop.::Do not copy heavy process." \
  --user-reference "A product or repo I admire::taste::Learn the clarity and operating standard.::Do not copy wording." \
  --local-constraint "Current Library Naming::structure::Keep naming aligned with the local skill library.::Do not inherit private references."
python3 scripts/yao.py skill-interpretation my-skill
python3 scripts/yao.py review-viewer my-skill
python3 scripts/yao.py review-studio my-skill
python3 scripts/yao.py artifact-design-profile my-skill
python3 scripts/yao.py prompt-quality-profile my-skill
python3 scripts/yao.py system-model my-skill
python3 scripts/yao.py feedback my-skill --note "Tighten exclusions before adding scripts." --rating 4 --category boundary
python3 scripts/yao.py adapt-scan my-skill --source ./curated-user-signals.jsonl
python3 scripts/yao.py adapt-propose my-skill
python3 scripts/yao.py daily-skillops my-skill --source ./curated-user-signals.jsonl
python3 scripts/yao.py weekly-curator my-skill
python3 scripts/yao.py adoption-drift my-skill --record-event skill_activation --activation-type explicit --outcome accepted
YAO_CLI_TELEMETRY=1 python3 scripts/yao.py validate my-skill
python3 scripts/yao.py telemetry-emit my-skill --event skill_activation --activation-type explicit --outcome accepted --command browser-extension
python3 scripts/yao.py telemetry-hooks my-skill
python3 scripts/telemetry_native_host.py my-skill --write-launcher /tmp/yao-telemetry-host.sh --write-manifest /tmp/yao-telemetry-host.json --allowed-origin chrome-extension://aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa/
python3 scripts/yao.py telemetry-import my-skill --input-jsonl /tmp/external-client-events.jsonl --command browser-extension
python3 scripts/yao.py review-waivers my-skill --add-waiver --gate-key trust-report --reviewer "Yao Team" --reason "Known warning accepted for this release with bounded follow-up." --expires-at 2026-09-30
python3 scripts/yao.py review-waivers my-skill --add-waiver --gate-key permission-gates --reviewer "Yao Team" --reason "Permission warning accepted only for this non-governed release window." --expires-at 2026-09-30
python3 scripts/yao.py review-annotations my-skill --add-annotation --gate-key output-lab --target-path reports/output_quality_scorecard.md --line 1 --body "Clarify recorded fixture vs model-executed evidence before release."
python3 scripts/yao.py baseline-compare --self
python3 scripts/yao.py check-update --notice --self
python3 scripts/yao.py self-update --self
python3 scripts/yao.py skill-ir . --output-json skill-ir/examples/yao-meta-skill.json --self
python3 scripts/yao.py compile-skill . --target openai --target claude --target generic --target vscode --self
python3 scripts/yao.py package . --platform generic --output-dir dist --self
python3 scripts/yao.py output-eval --self
python3 scripts/yao.py output-exec --self
python3 scripts/yao.py output-review --self
python3 scripts/yao.py conformance . --self
python3 scripts/yao.py trust . --self
python3 scripts/yao.py python-compat . --self
python3 scripts/yao.py runtime-permissions . --package-dir dist --self
python3 scripts/yao.py skill-atlas --workspace-root . --self
python3 scripts/yao.py registry-audit . --self
python3 scripts/yao.py package-verify . --package-dir dist --require-zip --self
python3 scripts/yao.py install-simulate . --package-dir dist --self
python3 scripts/yao.py upgrade-check . --previous-package-json registry/examples/yao-meta-skill-1.0.0.json --self
python3 scripts/yao.py world-class-evidence . --self
SUBMISSIONS_DIR="${SUBMISSIONS_DIR:-evidence/world_class/submissions}"
python3 scripts/yao.py world-class-preflight . --submissions-dir "$SUBMISSIONS_DIR" --self
python3 scripts/yao.py world-class-submission-kit . --output-dir "$SUBMISSIONS_DIR" --self
# Alternative: prefill artifact SHA-256 digests while keeping drafts template-only.
python3 scripts/yao.py world-class-submission-kit . --output-dir "$SUBMISSIONS_DIR" --prefill-artifacts --self
python3 scripts/yao.py world-class-intake . --submissions-dir "$SUBMISSIONS_DIR" --self
python3 scripts/yao.py world-class-submission-review . --submissions-dir "$SUBMISSIONS_DIR" --self
python3 scripts/yao.py world-class-ledger . --submissions-dir "$SUBMISSIONS_DIR" --self
python3 scripts/yao.py world-class-runbook . --submissions-dir "$SUBMISSIONS_DIR" --self
python3 scripts/yao.py world-class-claim-guard . --self
python3 scripts/yao.py benchmark-reproducibility . --self
python3 scripts/yao.py evidence-consistency . --self

Local Development Source

Development source: this repository is the source of truth for authoring and review.

Use Python 3.11 or newer for local development. GitHub Actions runs the test suite on Python 3.11, and the Makefile checks the active interpreter before running make test or make ci-test.

python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install --requirement requirements-ci.txt
make ci-test

If python3 points to an older system interpreter, pass the interpreter explicitly:

make PYTHON=python3.11 ci-test

Disabled mirror: ~/.agents/skills.disabled/yao-meta-skill is the local backup mirror for this source. Keeping the mirror outside ~/.agents/skills prevents Codex from showing a duplicate Yao Meta Skill while this repository is also visible in the active workspace.

Sync the current source into the disabled mirror:

make sync-local-install

The sync command first rebuilds the package and runs install preflight against dist/yao-meta-skill.zip. It refuses to sync when package extraction, adapter readability, or installer permission enforcement fails. After the preflight passes, it copies Git-tracked files plus new source files in code and guidance directories such as scripts/, tests/, references/, and docs/. It skips untracked business-skill folders and untracked private reports by default, so local experiments do not leak into the mirror.

Restore an active global Codex install only when you intentionally want this skill discoverable outside the development workspace:

make sync-active-install

That active install writes to ~/.agents/skills/yao-meta-skill and can make Codex show a second Yao Meta Skill entry while this repository is open as a skills workspace.

Generated Artifact Boundaries

Keep this repository focused on the meta-skill factory.

  • Put reusable factory examples in examples/.
  • Name embedded example entrypoints SKILL.example.md; reserve the exact SKILL.md filename for the installable root skill so recursive agent discovery does not activate examples or test fixtures.
  • Put reusable benchmark evidence, regression results, and release evidence in reports/.
  • Keep private analysis reports, customer-specific outputs, and one-off generated business skills outside this repository unless they are intentionally promoted into an example or regression fixture.
  • Place real generated skills as sibling skill directories under the local skill workspace, not as top-level folders inside yao-meta-skill.

5-Minute Workflow

  1. Start from a raw workflow note.
  2. Turn it into a skill package with SKILL.md, agents/interface.yaml, and only the folders the workflow actually needs.
  3. Validate the trigger description with evals/trigger_cases.json.
  4. Export compatibility artifacts for the clients you care about.
  5. Compare the result against the examples in examples/.

Minimum commands:

python3 scripts/trigger_eval.py --description-file evals/improved_description.txt --cases evals/trigger_cases.json
python3 scripts/run_description_optimization_suite.py
python3 scripts/judge_blind_eval.py --description-file SKILL.md --cases evals/blind_holdout/trigger_cases.json --semantic-config evals/semantic_config.json
python3 scripts/context_sizer.py .
python3 scripts/resource_boundary_check.py .
python3 scripts/governance_check.py . --require-manifest
python3 scripts/compile_skill.py .
python3 scripts/cross_packager.py . --platform openai --platform claude --platform generic --platform vscode --expectations evals/packaging_expectations.json --zip
python3 scripts/probe_runtime_permissions.py . --package-dir dist
python3 tests/verify_packager_failures.py

Or run everything together:

make test

Unified authoring flow:

python3 scripts/yao.py init my-skill --description "Describe what the skill does."
python3 scripts/yao.py validate my-skill
python3 scripts/yao.py workspace-flow --target root --label first-pass --self
python3 scripts/yao.py review-viewer my-skill
python3 scripts/yao.py review --target root --self
python3 scripts/yao.py release-snapshot --target root --label release-candidate --self
python3 scripts/yao.py skill-ir . --output-json skill-ir/examples/yao-meta-skill.json --self
python3 scripts/yao.py compile-skill . --self
python3 scripts/yao.py package . --platform openai --platform claude --platform generic --platform vscode --output-dir dist --zip --self
python3 scripts/yao.py runtime-permissions . --package-dir dist --self
python3 scripts/yao.py package-verify . --package-dir dist --require-zip --self
python3 scripts/yao.py test --self

Results

The homepage panel below is generated from the current eval suite so the family-level outcome is visible without opening raw JSON.

  • regression corpus: 66 prompts across 21 families
  • aggregate result: 0 false positives, 0 false negatives, average precision 1.0, average recall 1.0
  • suite status:
Suite Cases FP FN Precision Recall
train 31 0 0 1.0 1.0
dev 22 0 0 1.0 1.0
holdout 13 0 0 1.0 1.0
Family Cases Pass Rate
brainstorm_only 2 1.0
brainstorm_vs_build 1 1.0
complex_multi_asset 3 1.0
document_export_vs_agent_skill 4 1.0
document_only 3 1.0
explain_not_package 1 1.0
explain_only 5 1.0
future_outline_vs_build 4 1.0
iterate_existing_skill 5 1.0
long_context_document_only 3 1.0
long_context_near_neighbor 3 1.0
long_context_summary_only 2 1.0
long_context_trigger 4 1.0
meta_skill_creation 1 1.0
one_off_vs_reusable 2 1.0
package_for_team 2 1.0
paraphrase_trigger 5 1.0
partial_scaffold_not_full_skill 4 1.0
summary_only 3 1.0
translate_only 4 1.0
workflow_to_skill 5 1.0

Full reports: reports/eval_suite.json and reports/family_summary.md

Current Strengths

The latest weighted review puts Yao at 91.5/100. The strongest dimensions are the ones that matter most when skills become long-lived team assets:

  • Method depth 9.5: formal skill engineering doctrine, archetypes, gate selection, non-skill decisions, lifecycle governance, and resource boundaries.
  • Toolchain completeness 9.5: authoring, validation, benchmark scan, description optimization, report generation, promotion checks, packaging, CI, and portability checks are wired into one operational flow.
  • Eval and test rigor 9.5: trigger quality is checked with train/dev/holdout, blind holdout, adversarial holdout, judge-backed blind eval, route confusion, drift history, and promotion gates.
  • Governance and lifecycle 9.5: important skills can carry owner, lifecycle state, review cadence, maturity score, trust boundaries, promotion decisions, and regression history.
  • Local execution reliability 9.5: the repository is executable locally through make test, make ci-test, and the unified scripts/yao.py authoring CLI.
  • Portability and distribution 9.0: neutral source metadata, client adapters, degradation rules, packaging contracts, and portability scoring preserve reusable semantics across target environments.
  • Systems stability: generated skills now include a system model that turns boundary discipline, feedback loops, drift watch, and leverage-point analysis into reviewer-visible evidence.
  • Context discipline 8.0: the entrypoint is still held under budget, but this is tracked as a live constraint because the system now carries more reports, examples, benchmark assets, and generated evidence.
  • Onboarding and review experience 6.5: quickstart, HTML overview, side-by-side review viewer, and feedback logs have improved the first-run experience, but this remains the clearest UX improvement area.

The current direction is deliberate: keep the entrypoint light, make evaluation hard to fake, make governance visible, and continue reducing the friction of first-time creation and review.

Why Yao

  • Lightweight: the entrypoint stays compact, context budgets are explicit, and extra structure is added only when it pays for itself.
  • Rigorous: trigger quality is checked with family regressions, blind holdout, adversarial holdout, route confusion, judge-backed blind eval, and promotion gates.
  • Governed: important skills are treated as maintainable assets with lifecycle state, maturity expectations, ownership, and review cadence.
  • Portable: source metadata stays neutral while adapters, degradation rules, and packaging contracts preserve reusable semantics across environments.

What It Does

This project helps you create, refactor, evaluate, and package skills as durable capability bundles rather than one-off prompts.

The design logic is simple:

  1. Capture the real recurring job behind the user's request.
  2. Set a clean skill boundary so one package does one coherent job.
  3. Optimize the trigger description before over-writing the body.
  4. Keep the main skill file small and move details into references or scripts.
  5. Add quality gates only when they pay for themselves.
  6. Export compatibility artifacts only for the clients you actually need.

Method Doctrine

The repository now treats method as a first-class asset instead of scattered guidance.

Why It Exists

Most teams keep valuable operating knowledge scattered across chats, personal prompts, oral habits, and undocumented workflows. This project converts that hidden process knowledge into:

  • discoverable skill packages
  • repeatable execution flows
  • lower-context instructions
  • reusable team assets
  • compatibility-ready distributions

Repository Structure

yao-meta-skill/
├── SKILL.md
├── README.md
├── VERSION
├── LICENSE
├── .gitignore
├── agents/
│   └── interface.yaml
├── evals/
├── examples/
├── references/
├── scripts/
└── templates/

Core Components

SKILL.md

The main skill entrypoint. It defines the trigger surface, operating modes, compact workflow, and output contract.

agents/interface.yaml

The neutral metadata source of truth. It stores display and compatibility metadata without locking the source tree to one vendor-specific path.

references/

Long-form material that should not bloat the main skill file. This includes design rules, evaluation guidance, compatibility strategy, and quality rubrics.

scripts/

Utility scripts that make the meta-skill operational:

  • trigger_eval.py: evaluates trigger descriptions with semantic intent concepts, explicit exclusions, and near-neighbor prompts
  • run_eval_suite.py: runs train/dev/holdout trigger suites, reports family-level regressions, and fails if aggregate regressions appear
  • optimize_description.py: generates candidate descriptions, scores them on dev, visible holdout, blind holdout, and adversarial holdout suites, then reports calibration and family health
  • judge_blind_eval.py: applies an independent rubric judge to blind-holdout prompts so blind acceptance is not backed only by the main threshold scorer
  • run_description_optimization_suite.py: runs description optimization across the root skill and governed examples, then writes reusable reports and optional drift snapshots with calibration and family summaries
  • promotion_checker.py: applies promotion policy to current description candidates, writes promotion decisions, builds candidate registries, and emits iteration bundles with review stubs
  • create_iteration_snapshot.py: freezes the current promotion decision into a versioned release snapshot with review, route, and context evidence
  • yao.py: unified authoring CLI that exposes init, validate, optimize-description, promote-check, python-compat, review, release-snapshot, workspace-flow, report, skill-report, skill-interpretation, skill-ir, compile-skill, output-exec, output-review, skill-os2-audit, skill-os2-coverage, world-class-evidence, world-class-ledger, world-class-intake, world-class-preflight, world-class-submission-kit, world-class-submission-review, world-class-runbook, world-class-claim-guard, benchmark-reproducibility, evidence-consistency, adapt-scan, adapt-propose, adapt-apply, daily-skillops, weekly-curator, telemetry-emit, telemetry-hooks, telemetry-import, package, registry-audit, package-verify, install-simulate, upgrade-check, review-waivers, and test as one entrypoint
  • render_description_drift_history.py: turns description-optimization snapshots into a readable drift-history report
  • build_confusion_matrix.py: scores route confusion across tracked sibling skills and no_route cases, then writes a route scorecard and optional milestone snapshot
  • render_iteration_ledger.py: compresses regression milestones, description optimization drift, and route scorecards into one iteration-facing ledger
  • context_sizer.py: estimates context weight and warns when the initial load gets too large
  • resource_boundary_check.py: audits whether detail is split across SKILL.md, references/, scripts/, assets/, and evals/ appropriately
  • governance_check.py: validates owner, review cadence, lifecycle stage, and maturity metadata
  • render_context_reports.py: generates root and example context-budget reports plus a shared context summary
  • render_regression_history.py: turns milestone snapshots into a readable regression history report
  • render_skill_os2_audit.py: renders a requirement-by-requirement Skill OS 2.0 audit that separates landed local evidence from human-required and external-required gaps
  • render_skill_os2_coverage.py: maps the Skill OS 2.0 upgrade blueprint to local artifacts, commands, tests, and remaining evidence boundaries
  • render_daily_skillops_report.py: renders an explicit-source Daily SkillOps operations report that summarizes redacted user patterns, proposal-only adaptations, approval state, release evidence, and world-class evidence gaps without scanning private logs or applying patches
  • render_weekly_curator_report.py: renders a weekly SkillOps curator report from generated daily reports, Skill Atlas, benchmark lock, evidence consistency, and world-class ledger state without scanning private logs or applying patches
  • skillops_opportunity.py: scores redacted SkillOps opportunities and maps them to approval-gated action types such as report-only, AGENTS update, existing-skill patch, or eval addition
  • render_world_class_evidence_plan.py: renders executable evidence tasks for remaining world-class gaps without treating planned external work as completed evidence
  • render_world_class_evidence_ledger.py: renders a machine-checkable ledger for current world-class evidence acceptance, anti-overclaim guards, provenance requirements, and privacy contracts
  • render_world_class_evidence_intake.py: validates world-class external and human evidence packets against provenance, privacy, artifact, and anti-overclaim requirements before ledger review
  • render_world_class_preflight.py: renders redacted collection preflight checks for pending provider, human, native-permission, and native-client evidence without accepting evidence
  • render_world_class_submission_review.py: renders a read-only queue that compares submissions, intake validation, source evidence, and ledger state without accepting evidence
  • render_world_class_operator_runbook.py: renders an operator-facing checklist and command map for collecting pending world-class evidence without accepting evidence
  • render_world_class_claim_guard.py: scans README, docs, and reports for premature world-class completion claims while accepted evidence is still pending
  • render_benchmark_reproducibility.py: renders methodology, artifact, failure-disclosure, and reproduction-command evidence for public benchmark claims
  • render_evidence_consistency.py: compares generated report facts across benchmark reproducibility, overview, interpretation, adoption drift, world-class ledger, coverage, and Review Studio artifacts
  • python_compat_check.py: checks Python source for supported-runtime compatibility hazards such as Python 3.11 f-string expression backslashes
  • cross_packager.py: builds client-specific export artifacts from Skill IR plus neutral metadata, with explicit platform contracts and validation
  • render_portability_report.py: scores cross-environment portability from neutral metadata, degradation rules, and consumer validation coverage
  • render_skill_overview.py: generates the white-background bilingual HTML skill audit report with sticky four-character Chinese navigation, top-right language switch, v2 scorecard, inline SVG charts, contract boundary, quality review, risk governance, assets, and iteration roadmap
  • render_skill_interpretation.py: renders reports/skill-interpretation.html/json as the first-class post-creation interpretation report while reusing the Skill Overview v2 model and Kami white layout
  • export_skill_ir.py: exports the 2.0 platform-neutral Skill IR contract from SKILL.md, manifest, interface metadata, evals, resources, and reports
  • compile_skill.py: compiles Skill IR into target-specific semantic contracts, generated-file maps, adapter modes, target-native behavior contracts, preserved semantics, warnings, and unsupported-feature notes
  • run_output_eval.py: runs the Output Eval Lab v0 with static with-skill vs baseline assertion grading, blind A/B review pack generation, and separate answer key artifacts
  • run_output_execution.py: records output-eval execution evidence, distinguishing recorded fixtures, command runners, and provider-backed model runs with timing and token metadata
  • local_output_eval_runner.py: deterministic local runner for command-executed output-eval smoke evidence without claiming provider-backed model generation
  • adjudicate_output_review.py: records reviewer choices for blind A/B output evals, compares them with the answer key, and renders pending, match, disagreement, and invalid-decision audit reports
  • render_review_annotations.py: records reviewer annotations tied to Review Studio gates, source/report paths, and optional line numbers, with open blocker annotations reflected in Review Studio decisions
  • run_conformance_suite.py: verifies runtime conformance for OpenAI, Claude, Agent Skills, VS Code/Copilot-style, and generic targets
  • trust_check.py: generates the trust/security report for scripts, dependencies, secret risk, bounded network host policy, execution-level --help smoke checks, permission inputs, trust metadata, and stable source-contract integrity
  • build_skill_atlas.py: builds the Skill Atlas catalog, route-overlap matrix, dependency graph, stale report, owner gaps, aggregate drift signals, and HTML overview for a multi-skill workspace
  • registry_audit.py: builds registry package metadata and audits version, owner, license, checksum, Skill IR source, and compatibility matrix
  • verify_package.py: verifies generated package manifests, target adapters, zip archive safety, archive checksum, and registry parity
  • simulate_install.py: extracts a generated zip into a temporary skill root and verifies entrypoint, manifest, interface, reports, and adapters can be loaded
  • upgrade_check.py: compares current and previous registry package metadata, recommends a version bump, and blocks incompatible upgrade claims
  • render_adoption_drift_report.py: records metadata-only local telemetry and renders adoption, missed-trigger, bad-output, script-error, and review-drift signals without packaging raw event logs
  • import_telemetry_events.py: imports external metadata-only telemetry JSONL after whole-file privacy validation, then refreshes the aggregate adoption drift report
  • emit_telemetry_event.py: emits one metadata-only external client event into a local spool for later telemetry-import, with dry-run validation and raw-content field blocking
  • render_telemetry_hook_recipes.py: renders Browser, Chrome, VS Code, CLI wrapper, and provider-adapter telemetry hook recipes with dry-run commands and explicit native-integration caveats
  • telemetry_native_host.py: receives Browser/Chrome Native Messaging length-prefixed JSON events, rejects raw-content fields, appends metadata-only events, and writes local launcher/manifest files for operator installation
  • yao_cli_telemetry.py: opt-in metadata-only yao.py run capture for command name, source, outcome, and failure class without command arguments or raw content
  • render_review_waivers.py: validates human reviewer risk approvals with gate keys, reasons, expiry dates, and blocker-safe waiver policy
  • init_skill.py, lint_skill.py, validate_skill.py, diff_eval.py: minimal authoring toolchain
  • check_update.py: checks GitHub for a newer VERSION or remote manifest version, stores its cache in the user cache directory, and leaves Skill source/install directories unchanged
  • render_output_risk_profile.py: predicts output-specific failure modes such as generic headings, citation clutter, screenshot mistakes, weak Markdown tables, and missing execution assumptions

evals/

Reusable trigger and packaging checks, including baseline and improved descriptions for comparison plus the root semantic configuration that drives description optimization.

This directory also contains route confusion fixtures and promotion policy rules for deciding when a route is promotable.

examples/

End-to-end examples showing raw workflow input, design summary, final generated skill shape, and targeted description-optimization packs where route wording is tuned against example-specific dev and holdout cases.

.github/workflows/test.yml

Continuous integration entrypoint that runs the full local regression suite on push and pull request.

Validation Notes

  • Trigger evaluation now uses a local semantic-intent model with explicit positive concepts, exclusion concepts, and boundary-case reporting.
  • The sample trigger report now covers a larger positive, negative, and near-neighbor set rather than a tiny demo set.
  • Train/dev/holdout trigger suites now separate iterative tuning from final verification.
  • Description optimization now uses dev for ranking, visible holdout for non-regression, blind holdout for acceptance, and adversarial holdout for harder route-collision checks without feeding the ranking loop.
  • Judge-backed blind eval now adds a rubric-based second opinion for blind prompts, so blind acceptance is not decided by one scorer alone.
  • Description drift history now records adversarial calibration gaps and family coverage, so routing changes can be judged on confidence and family stability rather than raw error counts alone.
  • Route confusion is now tracked explicitly across the root meta-skill, frontend review skill, governed incident skill, and no_route cases, so route theft is visible instead of implicit.
  • Promotion policy now requires visible holdout, blind holdout, adversarial holdout, and route confusion to stay clean before a description should be considered promotable.
  • Promotion checking now emits explicit decisions, candidate lifecycle states, iteration bundles, and human-review stubs rather than leaving promotion as a prose-only step.
  • Promotion decisions now distinguish “no candidate beat current” from “current still has residual route risk,” so iteration can be audited without forcing every issue into a false block.
  • Packaging validation now uses explicit contracts and YAML parsing, but it is still a lightweight local validation layer rather than a full platform integration suite.
  • evals/failure-cases.md captures known weak spots that should remain part of regression checks.
  • failures/ captures reusable anti-pattern writeups and machine-runnable failure cases for routing, packaging, and authoring failures.
  • tests/verify_packager_failures.py checks that invalid metadata, invalid YAML, and unsupported targets fail clearly.
  • Governance metadata and resource-boundary rules now have runnable checks instead of staying as prose only.
  • Governance checks now emit a maturity score so governed assets can be compared instead of only pass/fail checked.
  • Description optimization drift history is now versioned separately from the main trigger regression history so routing improvements are visible over time.
  • Iteration evidence now records why a candidate was kept, blocked, or promotable via a shared regression-cause taxonomy and bundle artifacts.
  • Declared maturity tiers are checked against recommended minimum governance scores, so production, library, and governed assets can be compared without forcing every strong example into the same label.
  • Context budgets are now tiered and explicit, so a governed skill can still choose a stricter production-sized initial-load budget.
  • Resource-boundary checks now detect decorative directories and compute a local quality-density signal instead of only checking raw token counts.

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions
+115.1K-115.1KWeek of 2026-03-29: +48,697 linesWeek of 2026-03-29: -5,752 linesWeek of 2026-04-05: +4,353 linesWeek of 2026-04-05: -644 linesWeek of 2026-04-12: +1,250 linesWeek of 2026-04-12: -127 linesWeek of 2026-04-19: +2,597 linesWeek of 2026-04-19: -319 linesWeek of 2026-04-26: +2,755 linesWeek of 2026-04-26: -698 linesWeek of 2026-05-03: +1,303 linesWeek of 2026-05-03: -92 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +1,076 linesWeek of 2026-05-17: -55 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +60,806 linesWeek of 2026-06-07: -4,321 linesWeek of 2026-06-14: +115,084 linesWeek of 2026-06-14: -17,086 linesWeek of 2026-06-21: +1,320 linesWeek of 2026-06-21: -497 linesWeek of 2026-06-28: +8,627 linesWeek of 2026-06-28: -2,905 linesWeek of 2026-07-05: +70 linesWeek of 2026-07-05: -0 linesWeek of 2026-07-12: +2,071 linesWeek of 2026-07-12: -898 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesWeek of 2026-08-02: +0 linesWeek of 2026-08-02: -0 linesWeek of 2026-08-09: +21,086 linesWeek of 2026-08-09: -11,832 linesWeek of 2026-08-16: +24,904 linesWeek of 2026-08-16: -15,569 linesWeek of 2026-08-23: +0 linesWeek of 2026-08-23: -0 linesWeek of 2026-08-30: +0 linesWeek of 2026-08-30: -0 linesWeek of 2026-09-06: +0 linesWeek of 2026-09-06: -0 linesWeek of 2026-09-13: +0 linesWeek of 2026-09-13: -0 linesMar 29, 2026Sep 13, 2026
+296K lines added, -60.8K removed over the last year.

Commits per week

last 52 weeks
310Week of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 30 commitsWeek of 2026-04-05: 15 commitsWeek of 2026-04-12: 5 commitsWeek of 2026-04-19: 8 commitsWeek of 2026-04-26: 2 commitsWeek of 2026-05-03: 2 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 1 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 5 commitsWeek of 2026-06-14: 7 commitsWeek of 2026-06-21: 1 commitsWeek of 2026-06-28: 3 commitsWeek of 2026-07-05: 1 commitsWeek of 2026-07-12: 6 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsWeek of 2026-08-02: 0 commitsWeek of 2026-08-09: 21 commitsWeek of 2026-08-16: 31 commitsWeek of 2026-08-23: 0 commitsWeek of 2026-08-30: 0 commitsWeek of 2026-09-06: 0 commitsWeek of 2026-09-13: 0 commitsWeek of 2026-09-20: 0 commitsWeek of 2026-09-27: 0 commitsOct 5, 2025Sep 27, 2026
138 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 0 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 7 commitsSun 11:00 — 1 commitsSun 12:00 — 3 commitsSun 13:00 — 0 commitsSun 14:00 — 10 commitsSun 15:00 — 0 commitsSun 16:00 — 1 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 1 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 0 commitsMon 10:00 — 2 commitsMon 11:00 — 7 commitsMon 12:00 — 0 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 0 commitsMon 16:00 — 0 commitsMon 17:00 — 4 commitsMon 18:00 — 0 commitsMon 19:00 — 0 commitsMon 20:00 — 1 commitsMon 21:00 — 0 commitsMon 22:00 — 0 commitsMon 23:00 — 0 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 1 commitsTue 11:00 — 2 commitsTue 12:00 — 0 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 0 commitsTue 16:00 — 0 commitsTue 17:00 — 0 commitsTue 18:00 — 0 commitsTue 19:00 — 1 commitsTue 20:00 — 5 commitsTue 21:00 — 8 commitsTue 22:00 — 2 commitsTue 23:00 — 1 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 2 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 3 commitsWed 8:00 — 0 commitsWed 9:00 — 3 commitsWed 10:00 — 1 commitsWed 11:00 — 4 commitsWed 12:00 — 0 commitsWed 13:00 — 1 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 0 commitsWed 18:00 — 1 commitsWed 19:00 — 3 commitsWed 20:00 — 3 commitsWed 21:00 — 3 commitsWed 22:00 — 1 commitsWed 23:00 — 9 commitsThu 0:00 — 0 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 1 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 1 commitsThu 12:00 — 0 commitsThu 13:00 — 0 commitsThu 14:00 — 7 commitsThu 15:00 — 3 commitsThu 16:00 — 1 commitsThu 17:00 — 3 commitsThu 18:00 — 0 commitsThu 19:00 — 1 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 5 commitsThu 23:00 — 0 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 1 commitsFri 10:00 — 0 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 1 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 1 commitsFri 18:00 — 8 commitsFri 19:00 — 0 commitsFri 20:00 — 0 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 2 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 2 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 1 commitsSat 19:00 — 1 commitsSat 20:00 — 2 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits149 (99%)
Community commits2 (1%)

151 commits in total over the last year.

DateListRankStars gained
Jun 19, 2026daily#14+13
  • public-apis/public-apis

    A collective list of free APIs

    486.1K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    373.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    285.8K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript