Trending repositories: llm-evaluation

4 tracked repositories tagged with llm-evaluation, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

4 of 4 repositories

  • promptfoo/promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.

    AI summary: A testing and red-teaming framework for evaluating prompts, RAG pipelines, and LLM applications.

    23,975developer-toolsTypeScriptMIT
  • HKUDS/ClawWork

    "ClawWork: OpenClaw as Your AI Coworker - πŸ’° $15K earned in 11 Hours"

    AI summary: A real-world economic benchmark where AI agents complete professional tasks to earn income and survive.

    8,299ai-mlPythonMIT
  • ifixai-ai/iFixAi

    Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.

    AI summary: A fast, independent auditing CLI to evaluate AI agent safety, alignment, and hallucinations.

    6,558securityPythonApache-2.0
  • MDX-Tom/gpt-5.6-instruct

    A Codex jailbreak prompt and test pack for gpt-5.6-sol. ι’ˆε―Ή gpt-5.6 η³»εˆ—ηš„ Codex η ΄η”²ζη€Ίθ―δΈŽζ΅‹θ―•εŒ…γ€‚

    AI summary: A comprehensive jailbreak prompt and testing suite specifically targeting the gpt-5.6-sol model.

    4,816securityPythonMIT