Trending repositories: evaluation
3 tracked repositories tagged with evaluation, ordered by stars. Use the topic filters below to narrow further.
3 of 3 repositories
promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
AI summary: A testing and red-teaming framework for evaluating prompts, RAG pipelines, and LLM applications.
23,975developer-toolsTypeScriptMITcobusgreyling/loop-engineering
Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani and Boris Cherny). Includes loop-audit, loop-init, loop-cost.
AI summary: A framework for building robust agentic loops and evaluation pipelines.
9,942ai-mlJavaScriptMITyaojingang/yao-meta-skill
YAO = Yielding AI Outcomes. A rigorous engineering, evaluation, governance, and portability system for reusable agent skills.
AI summary: A meta-learning framework for training AI agents to rapidly acquire new skills.
2,338ai-mlPythonMIT