muratcankoylan/Agent-Skills-for-Context-EngineeringPublic

A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.

AI summary: A comprehensive collection of skills and principles for engineering context and operating loops in production AI agent systems.

Stars
17.6K
+30 today
Forks
1.5K
Watchers
97
Open issues
16
Open PRs
24
Contributors
~11
Commits
193
Branches
38

PythonMITCreated Dec 21, 2025Last push 4d agoLatest release v2.3.0+115 stars this week+138 this month

Star history

since Dec 21, 2025
05K10K15KDec 2025Mar 2026May 2026Aug 2026
17.6K stars as of Aug 7, 2026, tracked back to Dec 21, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulMonWedFri2025-08-02: 0 commits2025-08-03: 0 commits2025-08-04: 0 commits2025-08-05: 0 commits2025-08-06: 0 commits2025-08-07: 0 commits2025-08-08: 0 commits2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 2 commits2025-12-21: 5 commits2025-12-22: 5 commits2025-12-23: 3 commits2025-12-24: 10 commits2025-12-25: 7 commits2025-12-26: 14 commits2025-12-27: 3 commits2025-12-28: 7 commits2025-12-29: 3 commits2025-12-30: 4 commits2025-12-31: 9 commits2026-01-01: 0 commits2026-01-02: 1 commit2026-01-03: 0 commits2026-01-04: 2 commits2026-01-05: 1 commit2026-01-06: 0 commits2026-01-07: 2 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 3 commits2026-01-12: 1 commit2026-01-13: 1 commit2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 1 commit2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 1 commit2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 1 commit2026-02-25: 0 commits2026-02-26: 1 commit2026-02-27: 2 commits2026-02-28: 0 commits2026-03-01: 1 commit2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 1 commit2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 1 commit2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 4 commits2026-03-18: 3 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 2 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 1 commit2026-04-10: 0 commits2026-04-11: 1 commit2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 1 commit2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 9 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 4 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 1 commit2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 0 commits2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 0 commits2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 2 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 0 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 3 commits2026-07-09: 1 commit2026-07-10: 0 commits2026-07-11: 11 commits2026-07-12: 0 commits2026-07-13: 4 commits2026-07-14: 4 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 0 commits2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits
143 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    17,627 stars

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    4 trending appearances

What Agent-Skills-for-Context-Engineering does

This repository shifts the focus from simple prompt engineering to 'context engineering'—the holistic management of a language model's entire attention budget. It solves the problem of unreliable AI agents by providing actionable skills for curating system prompts, tool definitions, message histories, and retrieved documents. The project offers a structured approach to designing agent operating loops and evaluating their behavior across any platform. It defines context engineering as a distinct discipline essential for building production-grade agentic systems. The materials guide developers on how to dynamically manage the context window to prevent hallucination and maintain agent focus during complex tasks.

AI engineers, system architects, and developers building production-grade LLM applications who need to move beyond basic prompting. Requires a solid understanding of language models and basic agent architectures.

  • Context engineering curriculum: Defines and teaches the specific discipline of managing an LLM's limited attention budget.
  • Holistic context curation: Provides strategies for balancing system prompts, tool outputs, and document retrieval simultaneously.
  • Operating loop design: Details the architectural patterns necessary for creating reliable, autonomous agent execution cycles.
  • Platform agnostic skills: Principles taught are applicable across any underlying LLM or agent framework.
  • Evaluation frameworks: Includes methodologies for assessing and testing agent behavior in production environments.

Where teams use it

Designing reliable autonomous agents

AI engineers can apply these principles to build agents that maintain focus over long execution loops without losing critical context.

Optimizing tool use

Developers struggling with agents that misuse or ignore tools can learn how to structure tool definitions within the context window.

Scaling RAG systems

Data scientists can utilize the context management techniques to ensure retrieved documents don't overwhelm the model's reasoning capabilities.

Standardizing AI development

Engineering teams can adopt this framework to create shared standards for context curation across their entire AI product suite.

Getting started: git clone https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering.git

README

main branch

Agent Skills for Context Engineering

A comprehensive, open collection of Agent Skills focused on context engineering and harness engineering principles for building production-grade AI agent systems. These skills teach the art and science of curating context, designing agent operating loops, and evaluating agent behavior across any agent platform.

Ask DeepWiki

What is Context Engineering?

Context engineering is the discipline of managing the language model's context window. Unlike prompt engineering, which focuses on crafting effective instructions, context engineering addresses the holistic curation of all information that enters the model's limited attention budget: system prompts, tool definitions, retrieved documents, message history, and tool outputs.

The fundamental challenge is that context windows are constrained not by raw token capacity but by attention mechanics. As context length increases, models exhibit predictable degradation patterns: the "lost-in-the-middle" phenomenon, U-shaped attention curves, and attention scarcity. Effective context engineering means finding the smallest possible set of high-signal tokens that maximize the likelihood of desired outcomes.

Recognition

This repository is cited in academic research as foundational work on static skill architecture:

"While static skills are well-recognized [Anthropic, 2025b; Muratcan Koylan, 2025], MCE is among the first to dynamically evolve them, bridging manual skill engineering and autonomous self-improvement."

  1. Meta Context Engineering via Agentic Skill Evolution, Peking University State Key Laboratory of General Artificial Intelligence (2025)
  2. Agent Harness Engineering: A Survey, CMU, Yale, JHU, NEU, Tulane, UAB, OSU, Virginia Tech, and Amazon (2026)

Skills Overview

Foundational Skills

These skills establish the foundational understanding required for all subsequent context engineering work.

Skill Description
context-fundamentals Understand what context is, why it matters, and the anatomy of context in agent systems
context-degradation Recognize patterns of context failure: lost-in-middle, poisoning, distraction, and clash
context-compression Design and evaluate compression strategies for long-running sessions

Architectural Skills

These skills cover the patterns and structures for building effective agent systems.

Skill Description
multi-agent-patterns Master orchestrator, peer-to-peer, and hierarchical multi-agent architectures
long-horizon-prompting NEW Write pseudo-formal task briefs for long-running autonomous agents and parallel orchestrations: exact success predicates, non-counting outcomes, audit-gated return conditions, effort floors, and diversity policies, modeled on the published GPT-5.6 Sol Ultra Cycle Double Cover prompt
memory-systems Design short-term, long-term, and graph-based memory architectures
tool-design Build tools that agents can use effectively
filesystem-context Use filesystems for dynamic context discovery, tool output offloading, and plan persistence
hosted-agents NEW Build background coding agents with sandboxed VMs, pre-built images, multiplayer support, and multi-client interfaces

Operational Skills

These skills address the ongoing operation and optimization of agent systems.

Skill Description
context-optimization Apply compaction, masking, and caching strategies
latent-briefing Share task-relevant orchestrator state with workers via task-guided KV cache compaction when the worker runtime is controllable
evaluation Build evaluation frameworks for agent systems
advanced-evaluation Master LLM-as-a-Judge techniques: direct scoring, pairwise comparison, rubric generation, and bias mitigation
harness-engineering Design autonomous agent harnesses with locked metrics, durable logs, novelty gates, rollback, and human approval boundaries
self-improvement-loops NEW Build loops where the harness itself is the optimization target: RSI, meta-harness search, failure-driven self-edits, evolutionary scaffold search, and acceptance gates for self-modifying systems

Development Methodology

These skills cover the meta-level practices for building LLM-powered projects.

Skill Description
project-development Design and build LLM projects from ideation through deployment, including task-model fit analysis, pipeline architecture, and structured output design

Cognitive Architecture Skills

These skills cover formal cognitive modeling for rational agent systems.

Skill Description
bdi-mental-states NEW Transform external RDF context into agent mental states (beliefs, desires, intentions) using formal BDI ontology patterns for deliberative reasoning and explainability

Design Philosophy

Progressive Disclosure

Each skill is structured for efficient context use. At startup, agents load only skill names and descriptions. Full content loads only when a skill is activated for relevant tasks.

Platform Agnosticism

These skills focus on transferable principles rather than vendor-specific implementations. The patterns work across Claude Code, Cursor, and any agent platform that supports skills or allows custom instructions.

Conceptual Foundation with Practical Examples

Scripts and examples demonstrate concepts using Python pseudocode that works across environments without requiring specific dependency installations.

Usage

Usage with Claude Code

This repository is a Claude Code Plugin Marketplace containing context engineering skills that Claude automatically discovers and activates based on your task context.

Installation

Step 1: Add the Marketplace

Run this command in Claude Code to register this repository as a plugin source:

/plugin marketplace add muratcankoylan/Agent-Skills-for-Context-Engineering

Step 2: Install the Plugin

Option A - Browse and install:

  1. Select Browse and install plugins
  2. Select context-engineering-marketplace
  3. Select context-engineering
  4. Select Install now

Option B - Direct install via command:

/plugin install context-engineering@context-engineering-marketplace

This installs all 17 skills in a single plugin. Skills are activated automatically based on your task context.

Skill Activation Scenarios

Skill Activate When
context-fundamentals Establishing context-window mental models, planning agent architecture, or explaining how context components affect model behavior
context-degradation Diagnosing attention failures, context poisoning, lost-in-middle behavior, or degraded agent performance across long sessions
context-compression Preserving useful state while reducing conversation, tool-output, or trajectory size under context pressure
context-optimization Improving token efficiency, retrieval precision, prefix reuse, masking, partitioning, or budget allocation for agent systems
latent-briefing Sharing orchestrator trajectory with workers via task-guided KV cache compaction when the worker runtime is controllable and the models are compatible
multi-agent-patterns Choosing coordination patterns, isolating context across agents, designing handoffs, or evaluating whether parallel agents are justified
long-horizon-prompting Writing or evaluating the launch prompt for a long-running autonomous agent or parallel orchestration: success predicates, non-counting outcomes, persistence and stop rules, adversarial audit gates, and portfolio diversity policies
memory-systems Persisting cross-session knowledge, tracking entities over time, choosing memory frameworks, or designing retrieval and update semantics
tool-design Defining agent-tool contracts, consolidating tool surfaces, improving descriptions, or making tool errors actionable
filesystem-context Moving large or durable context into files, creating scratchpads, supporting just-in-time discovery, or coordinating agents through shared artifacts
hosted-agents Running coding agents in remote sandboxes, background environments, warm pools, or multiplayer agent infrastructure
evaluation Creating deterministic checks, rubrics, regression suites, production monitoring, or quality gates for agent behavior
advanced-evaluation Using LLM judges, pairwise comparison, calibration, bias mitigation, or human-aligned quality assessment
harness-engineering Designing autonomous loops with locked evaluators, editable surfaces, durable logs, novelty gates, rollback, and approval boundaries
self-improvement-loops Building loops that modify themselves: failure-driven harness self-edits, meta-harness search, evolutionary scaffold search, context mechanism evolution, and acceptance gates for self-modification
project-development Deciding whether an LLM is appropriate, shaping batch pipelines, creating staged artifacts, or estimating operational cost
bdi-mental-states Modeling beliefs, desires, intentions, rational action traces, or neuro-symbolic state transformations for agents
Screenshot 2025-12-26 at 12 34 47 PM

For Cursor, Codex, and Open Plugins

This repository ships as an Open Plugins plugin. Hosts discover skills from the repo-root skills/ directory (each subdirectory contains a SKILL.md file). The manifest lives at .plugin/plugin.json.

Cursor (recommended):

  1. Install from the Cursor Plugin Directory, or clone this repo and point Cursor at the plugin root.
  2. Cursor reads .plugin/plugin.json and discovers the repo-root skills/ directory through the Open Plugins manifest.
  3. For project-local manual installs, copy skill directories into .cursor/skills/. Do not rely on repository symlinks; they are fragile on Windows and in plugin packaging.

Codex / GitHub Copilot CLI / other Open Plugins hosts:

  1. Clone or add this repository as a plugin directory.
  2. The host reads .plugin/plugin.json and discovers all 17 skills under skills/.
  3. For project-local manual installs, copy skill directories into .codex/skills/ or the host's documented Agent Skills directory.

Using Individual Skills

Agent Skills require a directory layout, not a flat markdown file. Copy the skill folder into your project's skills directory:

# Example: add just the context-fundamentals skill to a Cursor project
mkdir -p .cursor/skills
cp -R skills/context-fundamentals .cursor/skills/

# Claude Code project-scoped install (same directory layout)
mkdir -p .claude/skills
cp -R skills/context-fundamentals .claude/skills/

# Codex project-scoped install
mkdir -p .codex/skills
cp -R skills/context-fundamentals .codex/skills/

# Generic Agent Skills repo-scoped install (Codex/OpenAI, Copilot CLI, Open Plugins hosts)
mkdir -p .agents/skills
cp -R skills/context-fundamentals .agents/skills/

Do not flatten SKILL.md into a single file at .claude/skills/context-fundamentals.md. That breaks relative references/ paths and violates the Agent Skills directory spec used by Cursor, Claude Code, and Codex.

Available skills: context-fundamentals, context-degradation, context-compression, context-optimization, latent-briefing, multi-agent-patterns, long-horizon-prompting, memory-systems, tool-design, filesystem-context, hosted-agents, evaluation, advanced-evaluation, harness-engineering, self-improvement-loops, project-development, bdi-mental-states

For Custom Implementations

Extract the principles and patterns from any skill and implement them in your agent framework. The skills are deliberately platform-agnostic.

Examples

The examples folder contains complete system designs that demonstrate how multiple skills work together in practice.

Example Description Skills Applied
digital-brain-skill NEW Personal operating system for founders and creators. Complete Claude Code skill with 6 modules, 4 automation scripts context-fundamentals, context-optimization, memory-systems, tool-design, multi-agent-patterns, evaluation, project-development
x-to-book-system Multi-agent system that monitors X accounts and generates daily synthesized books multi-agent-patterns, memory-systems, context-optimization, tool-design, evaluation
llm-as-judge-skills Production-ready LLM evaluation tools with TypeScript implementation, 19 passing tests advanced-evaluation, tool-design, context-fundamentals, evaluation
book-sft-pipeline Train models to write in any author's style. Includes Gertrude Stein case study with 70% human score on Pangram, $2 total cost project-development, context-compression, multi-agent-patterns, evaluation
interleaved-thinking Reasoning trace optimizer that captures, analyzes, and converts agent failure patterns into generated skills evaluation, advanced-evaluation, context-degradation, harness-engineering
long-horizon-prompt-lab Production-ready educational website: method guide, copyable task-brief template, four complete prompt rewrites, structural audits, and a caveated research/vendor reference catalog long-horizon-prompting, harness-engineering, multi-agent-patterns, advanced-evaluation

Each example includes:

  • Complete PRD with architecture decisions
  • Skills mapping showing which concepts informed each decision
  • Implementation guidance

Digital Brain Skill Example

The digital-brain-skill example is a complete personal operating system demonstrating comprehensive skills application:

  • Progressive Disclosure: 3-level loading (SKILL.md → MODULE.md → data files)
  • Module Isolation: 6 independent modules (identity, content, knowledge, network, operations, agents)
  • Append-Only Memory: JSONL files with schema-first lines for agent-friendly parsing
  • Automation Scripts: 4 consolidated tools (weekly_review, content_ideas, stale_contacts, idea_to_draft)

Includes detailed traceability in HOW-SKILLS-BUILT-THIS.md mapping every architectural decision to specific skill principles.

LLM-as-Judge Skills Example

The llm-as-judge-skills example is a complete TypeScript implementation demonstrating:

  • Direct Scoring: Evaluate responses against weighted criteria with rubric support
  • Pairwise Comparison: Compare responses with position bias mitigation
  • Rubric Generation: Create domain-specific evaluation standards
  • EvaluatorAgent: High-level agent combining all evaluation capabilities

Book SFT Pipeline Example

The book-sft-pipeline example demonstrates training small models (8B) to write in any author's style:

  • Intelligent Segmentation: Two-tier chunking with overlap for maximum training examples
  • Prompt Diversity: 15+ templates to prevent memorization and force style learning
  • Tinker Integration: Complete LoRA training workflow with $2 total cost
  • Validation Methodology: Modern scenario testing proves style transfer vs content memorization

Integrates with context engineering skills: project-development, context-compression, multi-agent-patterns, evaluation.

Researcher Operating System

The researcher directory is a file-based operating system for turning external research into skill changes. It exists so this repository can act as a compounding source of truth instead of an anthology.

Measured router-benchmark results

The skill router (which decides whether the right skill gets loaded for a given task) has been benchmarked end-to-end against four frontier models via the Cursor SDK. Three full sweeps (50 prompts x 4 models x 3 replications = 600 calls each):

Per-skill effect size for the three skills the data flagged:

Skill Baseline top-1 After rewrite Delta
context-fundamentals 0.255 0.489 +23.4pp
project-development 0.750 1.000 +25pp (now perfect)
tool-design 0.729 0.807 +7.8pp

Per-model top-1 accuracy after the corpus-wide hardening pass:

Model Top-1 Top-3
gemini-3.1-pro 0.920 0.933
composer-2 0.913 0.947
gpt-5.5 0.913 0.973
claude-opus-4-7 0.840 0.933

Reproduce any of these numbers exactly via the runner under researcher/benchmarks/sdk-runner/.

What it includes

  • Source registry (researcher/source-registry.md): priority sources, exclusion rules, monitoring queries.
  • Rubrics (researcher/rubrics/): content curation, skill change, harness change, pairwise skill revision.
  • Mechanism registry (researcher/mechanisms/registry.jsonl + ledgers/): 16 accepted behavior changes used as the primary novelty signal, with append-only accepted/rejected ledgers for institutional memory.
  • Claim provenance (researcher/claims/index.jsonl): 12 provenance-tracked claims with source URL, evidence strength, volatility, and last reviewed date.
  • Corpus index (researcher/corpus/index.json): canonical machine-readable map of skills, activation scenarios, mechanism IDs, and claim IDs.
  • Run state machine (researcher/runs/<run-id>/run-state.json): initialized -> retrieved -> evaluated -> proposed -> novelty_checked -> validated -> pr_ready -> closed.
  • Activation regression tests (researcher/fixtures/activation-cases.jsonl): 19 deterministic prompts that catch skill-boundary confusion.
  • Adversarial benchmark harness (researcher/benchmarks/): scenarios that try to game the loop (duplicate mechanisms, unretrieved evidence, wrong rubric math, self-approved rubric changes, weak-evidence novelty).
  • Continuous loop (researcher/scripts/loop_*.py + researcher/orchestration/launchd/): inbox, source discovery, one-state-at-a-time advancement, daily ops, parked review queue, launchd service definitions.
  • Skill health gate (researcher/scripts/skill_health.py): deterministic body-quality scoring; current strict corpus score is 0.9117 with 0 flagged skills.

Operator commands

Install the validation dependencies once before running local gates:

python3 -m pip install -r requirements-dev.txt
# Deterministic gates (also run in CI on every PR)
python3 -m unittest researcher.scripts.tests.test_skill_frontmatter
python3 researcher/scripts/validate_platform_compat.py --require-reference-validator
python3 researcher/scripts/validate_repo.py --strict
python3 researcher/scripts/skill_health.py --strict --no-history
python3 researcher/scripts/run_benchmarks.py
python3 researcher/scripts/check_activation_cases.py

# Per-run readiness (active runs only)
python3 researcher/scripts/validate_run.py --run-dir researcher/runs/<run-id>

# Continuous loop, manual
python3 researcher/scripts/loop_discover.py
python3 researcher/scripts/loop_step.py --allow-fetch
python3 researcher/scripts/loop_daily.py
python3 researcher/scripts/loop_status.py

# Continuous loop, daemon (macOS)
researcher/orchestration/launchd/install.sh    # install launchd jobs (10-min step, 12h discover, daily ops)
researcher/orchestration/launchd/uninstall.sh  # remove launchd jobs

See researcher/runbooks/continuous-operation.md for daemon details, budgets, and the human review surface.

Guarantees

  • The loop never invokes paid LLMs or makes outbound writes; HTTP retrieval is stdlib-only with a 1.5 MB cap and a 30-second timeout.
  • Mechanism promotion requires a recorded human reviewer and a passing run-readiness check.
  • All queue mutations are atomic (temp file + os.replace) and serialized via fcntl locks.
  • Agents may prepare PRs after gates pass; merge and push remain human-controlled.

Star History

star-history-2026526

Structure

Each skill follows the Agent Skills specification:

skill-name/
├── SKILL.md              # Required: instructions + metadata
├── scripts/              # Optional: executable code demonstrating concepts
└── references/           # Optional: additional documentation and resources

See the template folder for the canonical skill structure.

Contributing

This repository follows the Agent Skills open development model. Contributions are welcome from the broader ecosystem. When contributing:

  1. Follow the skill template structure
  2. Provide clear, actionable instructions
  3. Include working examples where appropriate
  4. Document trade-offs and potential issues
  5. Keep SKILL.md under 500 lines for optimal performance

Feel free to contact Muratcan Koylan for collaboration opportunities or any inquiries.

License

MIT License - see LICENSE file for details.

References

The principles in these skills are derived from research and production experience at leading AI labs and framework developers. Each skill includes references to the underlying research and case studies that inform its recommendations.

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Releases and announcements

3 total
  1. The first release with measured benchmark results across four frontier models via the Cursor SDK, a corpus-wide hardening pass across all 15 published skills, and the first full-body Stage 3 effectiveness result. The correction in this release is important: we did not stop at the three descriptions that the router benchmark complained about. Every skill body was audited against the same standard, and the machine-readable substrate moved with it: mechanisms, claims, corpus index, activation fixtures, validators, template, docs, PR/release narrative. > Pull request: [#87](https://github.com/muratcankoylan/Agent-Skills-for-Context-Engineering/pull/87) > Full inventory: `CHANGELOG.md` > Latest router report: `researcher/benchmarks/router/results-published/2026-05-19.md` > Stage 3 hard pilot: `researcher/benchmarks/effectiveness/results-published/2026-05-19-stage3-hard-pilot.md` > Project narrative: `researcher/insights/how-we-built-this.md` > Technical findings: `researcher/insights/auto-research-experiment.md` > Benchmark plan: `researcher/benchmarks/PLAN.md` ## Research charts ![Research-to-Skill Operating System](https://raw.githubusercontent.com/muratcankoylan/Age

  2. ## What's Changed All 13 skills have been comprehensively rewritten based on learnings from Anthropic's ["Lessons from Building Claude Code: How We Use Skills"](https://www.anthropic.com) article. This is the largest single update to the skills collection. ### The Core Transformation Skills have been rewritten from **textbook voice** (explains concepts) to **hybrid instructional voice** (leads with actions, weaves in reasoning): **Before:** > "System prompts establish the agent's core identity, constraints, and behavioral guidelines. They are loaded once at session start and typically persist throughout the conversation." **After:** > "Organize system prompts into distinct sections using XML tags or Markdown headers. System prompts persist throughout the conversation, so place the most critical constraints at the beginning and end where attention is strongest — the middle receives 10-40% less recall accuracy." ### Changes Across All 13 Skills - **Hybrid voice rewrite** — every prose section rewritten from "X is Y" to "Do X because Y" while preserving all substantive knowledge (metrics, research findings, thresholds) - **Gotchas sections** — standardized `## Gotchas` added to

  3. ## 🎉 What's New ### New Skill: Advanced Evaluation A comprehensive skill for mastering LLM-as-a-Judge evaluation techniques. Based on research from [Eugene Yan's LLM-Evaluators](https://eugeneyan.com/writing/llm-evaluators/). **Covers:** - Direct scoring vs. pairwise comparison selection - Position, length, and verbosity bias mitigation - Metric selection (Cohen's κ, Spearman's ρ, Kendall's τ) - Production evaluation pipeline design - 10 actionable guidelines for reliable evaluation 📁 [`skills/advanced-evaluation/`](skills/advanced-evaluation/) ### New Example: LLM-as-Judge Skills A complete TypeScript [AI SDK-6](https://vercel.com/blog/ai-sdk-6) implementation demonstrating the Advanced Evaluation skill in practice. **Includes:** - 3 evaluation tools: `directScore`, `pairwiseCompare`, `generateRubric` - `EvaluatorAgent` class with full evaluation workflows - 19 passing tests with real OpenAI API calls - Position bias mitigation with automatic position swapping - Zod schemas for type-safe inputs/outputs 📁 [`examples/llm-as-judge-skills/`](examples/llm-as-judge-skills/) ## Quick Start cd examples/llm-as-judge-skills npm install cp env.example

Commits per week

last 52 weeks
470Week of 2025-08-02: 0 commitsWeek of 2025-08-09: 0 commitsWeek of 2025-08-16: 0 commitsWeek of 2025-08-23: 0 commitsWeek of 2025-08-30: 0 commitsWeek of 2025-09-06: 0 commitsWeek of 2025-09-13: 0 commitsWeek of 2025-09-20: 0 commitsWeek of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 2 commitsWeek of 2025-12-21: 47 commitsWeek of 2025-12-28: 24 commitsWeek of 2026-01-04: 5 commitsWeek of 2026-01-11: 5 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 1 commitsWeek of 2026-02-08: 1 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 4 commitsWeek of 2026-03-01: 2 commitsWeek of 2026-03-08: 1 commitsWeek of 2026-03-15: 7 commitsWeek of 2026-03-22: 2 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 2 commitsWeek of 2026-04-12: 1 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 0 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 9 commitsWeek of 2026-05-17: 4 commitsWeek of 2026-05-24: 1 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 0 commitsWeek of 2026-06-14: 0 commitsWeek of 2026-06-21: 0 commitsWeek of 2026-06-28: 2 commitsWeek of 2026-07-05: 15 commitsWeek of 2026-07-12: 8 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 0 commitsAug 2, 2025Jul 26, 2026
143 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 1 commitsSun 1:00 — 0 commitsSun 2:00 — 1 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 1 commitsSun 12:00 — 2 commitsSun 13:00 — 3 commitsSun 14:00 — 0 commitsSun 15:00 — 1 commitsSun 16:00 — 2 commitsSun 17:00 — 0 commitsSun 18:00 — 4 commitsSun 19:00 — 1 commitsSun 20:00 — 0 commitsSun 21:00 — 4 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 2 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 2 commitsMon 10:00 — 2 commitsMon 11:00 — 2 commitsMon 12:00 — 1 commitsMon 13:00 — 3 commitsMon 14:00 — 0 commitsMon 15:00 — 0 commitsMon 16:00 — 1 commitsMon 17:00 — 0 commitsMon 18:00 — 1 commitsMon 19:00 — 0 commitsMon 20:00 — 0 commitsMon 21:00 — 1 commitsMon 22:00 — 0 commitsMon 23:00 — 3 commitsTue 0:00 — 1 commitsTue 1:00 — 6 commitsTue 2:00 — 2 commitsTue 3:00 — 0 commitsTue 4:00 — 1 commitsTue 5:00 — 3 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 2 commitsTue 12:00 — 2 commitsTue 13:00 — 1 commitsTue 14:00 — 0 commitsTue 15:00 — 1 commitsTue 16:00 — 0 commitsTue 17:00 — 1 commitsTue 18:00 — 2 commitsTue 19:00 — 0 commitsTue 20:00 — 1 commitsTue 21:00 — 1 commitsTue 22:00 — 0 commitsTue 23:00 — 2 commitsWed 0:00 — 13 commitsWed 1:00 — 2 commitsWed 2:00 — 0 commitsWed 3:00 — 1 commitsWed 4:00 — 3 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 0 commitsWed 11:00 — 1 commitsWed 12:00 — 6 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 0 commitsWed 18:00 — 0 commitsWed 19:00 — 1 commitsWed 20:00 — 0 commitsWed 21:00 — 0 commitsWed 22:00 — 0 commitsWed 23:00 — 0 commitsThu 0:00 — 4 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 0 commitsThu 4:00 — 3 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 1 commitsThu 11:00 — 0 commitsThu 12:00 — 1 commitsThu 13:00 — 0 commitsThu 14:00 — 0 commitsThu 15:00 — 0 commitsThu 16:00 — 1 commitsThu 17:00 — 0 commitsThu 18:00 — 0 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 0 commitsThu 23:00 — 0 commitsFri 0:00 — 0 commitsFri 1:00 — 1 commitsFri 2:00 — 2 commitsFri 3:00 — 0 commitsFri 4:00 — 1 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 1 commitsFri 8:00 — 0 commitsFri 9:00 — 2 commitsFri 10:00 — 5 commitsFri 11:00 — 3 commitsFri 12:00 — 6 commitsFri 13:00 — 1 commitsFri 14:00 — 0 commitsFri 15:00 — 1 commitsFri 16:00 — 0 commitsFri 17:00 — 0 commitsFri 18:00 — 0 commitsFri 19:00 — 0 commitsFri 20:00 — 0 commitsFri 21:00 — 4 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 0 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 2 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 0 commitsSat 15:00 — 1 commitsSat 16:00 — 0 commitsSat 17:00 — 2 commitsSat 18:00 — 1 commitsSat 19:00 — 0 commitsSat 20:00 — 1 commitsSat 21:00 — 2 commitsSat 22:00 — 0 commitsSat 23:00 — 8 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits158 (82%)
Community commits35 (18%)

193 commits in total over the last year.

DateListRankStars gained
Feb 27, 2026daily#15+177
Feb 26, 2026daily#11+278
Feb 25, 2026daily#18+191
Feb 24, 2026daily#11+202