Trending repositories: evaluation

3 tracked repositories tagged with evaluation, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

3 of 3 repositories

  • promptfoo/promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

    AI summary: A testing and red-teaming framework for evaluating prompts, RAG pipelines, and LLM applications.

    23,975developer-toolsTypeScriptMIT
  • cobusgreyling/loop-engineering

    Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani and Boris Cherny). Includes loop-audit, loop-init, loop-cost.

    AI summary: A framework for building robust agentic loops and evaluation pipelines.

    9,942ai-mlJavaScriptMIT
  • yaojingang/yao-meta-skill

    YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.

    AI summary: A meta-learning framework for training AI agents to rapidly acquire new skills.

    2,338ai-mlPythonMIT