humanlayer/12-factor-agentsPublic

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

AI summary: Principles for building production-ready, reliable LLM-powered software and AI agents.

Stars
26.5K
+83 today
Forks
2K
Watchers
222
Open issues
15
Open PRs
12
Contributors
~16
Commits
273
Branches
16

TypeScriptOtherCreated Mar 30, 2025Last push 1y ago+166 stars this week+884 this month

Quick answers

What is 12-factor-agents?
Principles for building production-ready, reliable LLM-powered software and AI agents.
What does 12-factor-agents do?
This repository outlines the '12-factor-agents' methodology, a set of principles designed to guide the development of robust and reliable LLM-powered applications. Inspired by the original 12-factor app methodology, it adapts these concepts to the unique challenges of building AI agents. It addresses critical issues such as state management, observability, tool use, and evaluation in the context of non-deterministic models. The goal is to provide a framework that elevates AI software from experimental prototypes to production-grade systems. The guidelines provide a rigorous intellectual framework for transitioning AI prototypes into highly dependable enterprise architecture. It aggressively targets the most common failure modes of LLM applications, such as unbounded memory growth and state corruption.
Who is 12-factor-agents for?
This methodology is essential for software architects, backend engineers, and AI developers tasked with building and maintaining production-ready LLM applications.
How do I get started with 12-factor-agents?
Read the principles outlined in the repository's documentation.
How popular is 12-factor-agents on GitHub?
humanlayer/12-factor-agents has 26,530 stars and 1,984 forks on GitHub, and gained 166 stars in the last 7 days.
What license does 12-factor-agents use?
humanlayer/12-factor-agents is released under the Other license.

Star history

since Jul 29, 2026
010K20KJul 2026Aug 2026Sep 2026Oct 2026
26.5K stars as of Oct 2, 2026. Measured daily since Jul 29, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks

Signals and awards

derived from tracked data
  • Widely adopted

    26,530 stars

What 12-factor-agents does

This repository outlines the '12-factor-agents' methodology, a set of principles designed to guide the development of robust and reliable LLM-powered applications. Inspired by the original 12-factor app methodology, it adapts these concepts to the unique challenges of building AI agents. It addresses critical issues such as state management, observability, tool use, and evaluation in the context of non-deterministic models. The goal is to provide a framework that elevates AI software from experimental prototypes to production-grade systems. The guidelines provide a rigorous intellectual framework for transitioning AI prototypes into highly dependable enterprise architecture. It aggressively targets the most common failure modes of LLM applications, such as unbounded memory growth and state corruption.

This methodology is essential for software architects, backend engineers, and AI developers tasked with building and maintaining production-ready LLM applications.

  • Architectural Principles: Defines 12 core guidelines tailored for LLM and agent development.
  • Production Focus: Emphasizes reliability, scalability, and maintainability in real-world deployments.
  • State Management Strategies: Offers patterns for handling the complex state of long-running agent interactions.
  • Observability Guidelines: Details approaches for logging, tracing, and monitoring non-deterministic AI outputs.
  • Evaluation Frameworks: Discusses methods for systematically testing and evaluating agent performance.
  • Resilience Engineering: Focuses on building agents that degrade gracefully when encountering edge cases
  • Unbounded State Mitigation: Provides actionable patterns for handling infinite context growth in chat applications

Where teams use it

Designing New AI Systems

Architects use these principles as a blueprint when designing the architecture of a new LLM-powered application.

Refactoring Prototypes

Engineering teams apply these guidelines to transition a fragile AI prototype into a robust production service.

Establishing Best Practices

Organizations adopt this methodology to standardize how they build and deploy AI agents across different teams.

Evaluating Agent Reliability

Developers use the principles to audit existing AI systems and identify areas requiring improved observability or state management.

System Architecture Audit

Reviewing an existing multi-agent pipeline against the 12 factors to identify structural flaws.

Enterprise AI Deployment

Structuring a new product launch to ensure compliance with modern AI engineering standards.

Getting started: Read the principles outlined in the repository's documentation.

README

main branch

12-Factor Agents - Principles for building reliable LLM applications

In the spirit of 12 Factor Apps. The source for this project is public at https://github.com/humanlayer/12-factor-agents, and I welcome your feedback and contributions. Let's figure this out together!

Tip

Missed the AI Engineer World's Fair? Catch the talk here

Looking for Context Engineering? Jump straight to factor 3

Want to contribute to npx/uvx create-12-factor-agent - check out the discussion thread

Screenshot 2025-04-03 at 2 49 07 PM

Hi, I'm Dex. I've been hacking on AI agents for a while.

I've tried every agent framework out there, from the plug-and-play crew/langchains to the "minimalist" smolagents of the world to the "production grade" langraph, griptape, etc.

I've talked to a lot of really strong founders, in and out of YC, who are all building really impressive things with AI. Most of them are rolling the stack themselves. I don't see a lot of frameworks in production customer-facing agents.

I've been surprised to find that most of the products out there billing themselves as "AI Agents" are not all that agentic. A lot of them are mostly deterministic code, with LLM steps sprinkled in at just the right points to make the experience truly magical.

Agents, at least the good ones, don't follow the "here's your prompt, here's a bag of tools, loop until you hit the goal" pattern. Rather, they are comprised of mostly just software.

So, I set out to answer:

What are the principles we can use to build LLM-powered software that is actually good enough to put in the hands of production customers?

Welcome to 12-factor agents. As every Chicago mayor since Daley has consistently plastered all over the city's major airports, we're glad you're here.

Special thanks to @iantbutler01, @tnm, @hellovai, @stantonk, @balanceiskey, @AdjectiveAllison, @pfbyjy, @a-churchill, and the SF MLOps community for early feedback on this guide.

The Short Version: The 12 Factors

Even if LLMs continue to get exponentially more powerful, there will be core engineering techniques that make LLM-powered software more reliable, more scalable, and easier to maintain.

Visual Nav

factor 1 factor 2 factor 3
factor 4 factor 5 factor 6
factor 7 factor 8 factor 9
factor 10 factor 11 factor 12

How we got here

For a deeper dive on my agent journey and what led us here, check out A Brief History of Software - a quick summary here:

The promise of agents

We're gonna talk a lot about Directed Graphs (DGs) and their Acyclic friends, DAGs. I'll start by pointing out that...well...software is a directed graph. There's a reason we used to represent programs as flow charts.

010-software-dag

From code to DAGs

Around 20 years ago, we started to see DAG orchestrators become popular. We're talking classics like Airflow, Prefect, some predecessors, and some newer ones like (dagster, inggest, windmill). These followed the same graph pattern, with the added benefit of observability, modularity, retries, administration, etc.

015-dag-orchestrators

The promise of agents

I'm not the first person to say this, but my biggest takeaway when I started learning about agents, was that you get to throw the DAG away. Instead of software engineers coding each step and edge case, you can give the agent a goal and a set of transitions:

025-agent-dag

And let the LLM make decisions in real time to figure out the path

026-agent-dag-lines

The promise here is that you write less software, you just give the LLM the "edges" of the graph and let it figure out the nodes. You can recover from errors, you can write less code, and you may find that LLMs find novel solutions to problems.

Agents as loops

As we'll see later, it turns out this doesn't quite work.

Let's dive one step deeper - with agents you've got this loop consisting of 3 steps:

  1. LLM determines the next step in the workflow, outputting structured json ("tool calling")
  2. Deterministic code executes the tool call
  3. The result is appended to the context window
  4. Repeat until the next step is determined to be "done"
initial_event = {"message": "..."}
context = [initial_event]
while True:
  next_step = await llm.determine_next_step(context)
  context.append(next_step)

  if (next_step.intent === "done"):
    return next_step.final_answer

  result = await execute_step(next_step)
  context.append(result)

Our initial context is just the starting event (maybe a user message, maybe a cron fired, maybe a webhook, etc), and we ask the llm to choose the next step (tool) or to determine that we're done.

Here's a multi-step example:

027-agent-loop-animation.mp4
GIF Version

027-agent-loop-animation

Why 12-factor agents?

At the end of the day, this approach just doesn't work as well as we want it to.

In building HumanLayer, I've talked to at least 100 SaaS builders (mostly technical founders) looking to make their existing product more agentic. The journey usually goes something like:

  1. Decide you want to build an agent
  2. Product design, UX mapping, what problems to solve
  3. Want to move fast, so grab $FRAMEWORK and get to building
  4. Get to 70-80% quality bar
  5. Realize that 80% isn't good enough for most customer-facing features
  6. Realize that getting past 80% requires reverse-engineering the framework, prompts, flow, etc.
  7. Start over from scratch
Random Disclaimers

DISCLAIMER: I'm not sure the exact right place to say this, but here seems as good as any: this in BY NO MEANS meant to be a dig on either the many frameworks out there, or the pretty dang smart people who work on them. They enable incredible things and have accelerated the AI ecosystem.

I hope that one outcome of this post is that agent framework builders can learn from the journeys of myself and others, and make frameworks even better.

Especially for builders who want to move fast but need deep control.

DISCLAIMER 2: I'm not going to talk about MCP. I'm sure you can see where it fits in.

DISCLAIMER 3: I'm using mostly typescript, for reasons but all this stuff works in python or any other language you prefer.

Anyways back to the thing...

Design Patterns for great LLM applications

After digging through hundreds of AI libriaries and working with dozens of founders, my instinct is this:

  1. There are some core things that make agents great
  2. Going all in on a framework and building what is essentially a greenfield rewrite may be counter-productive
  3. There are some core principles that make agents great, and you will get most/all of them if you pull in a framework
  4. BUT, the fastest way I've seen for builders to get high-quality AI software in the hands of customers is to take small, modular concepts from agent building, and incorporate them into their existing product
  5. These modular concepts from agents can be defined and applied by most skilled software engineers, even if they don't have an AI background
The fastest way I've seen for builders to get good AI software in the hands of customers is to take small, modular concepts from agent building, and incorporate them into their existing product

The 12 Factors (again)

Honorable Mentions / other advice

Related Resources

Contributors

Thanks to everyone who has contributed to 12-factor agents!

dexhorthy Sypherd tofaramususa a-churchill Elijas hugolmn jeremypeters

kndl maciejkos pfbyjy 0xRaduan zyuanlim lombardo-chcg sahanatvessel

License

All content and images are licensed under a CC BY-SA 4.0 License

Code is licensed under the Apache 2.0 License

View on GitHub

Recent activity

commits and pull requests

Code frequency

additions and deletions

Commits per week

last 52 weeks

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 1 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 1 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 1 commitsSun 14:00 — 0 commitsSun 15:00 — 5 commitsSun 16:00 — 0 commitsSun 17:00 — 1 commitsSun 18:00 — 1 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 1 commitsSun 22:00 — 0 commitsSun 23:00 — 0 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 2 commitsMon 7:00 — 1 commitsMon 8:00 — 3 commitsMon 9:00 — 3 commitsMon 10:00 — 7 commitsMon 11:00 — 2 commitsMon 12:00 — 5 commitsMon 13:00 — 8 commitsMon 14:00 — 2 commitsMon 15:00 — 2 commitsMon 16:00 — 3 commitsMon 17:00 — 6 commitsMon 18:00 — 1 commitsMon 19:00 — 0 commitsMon 20:00 — 9 commitsMon 21:00 — 3 commitsMon 22:00 — 0 commitsMon 23:00 — 1 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 0 commitsTue 8:00 — 1 commitsTue 9:00 — 2 commitsTue 10:00 — 3 commitsTue 11:00 — 0 commitsTue 12:00 — 0 commitsTue 13:00 — 1 commitsTue 14:00 — 0 commitsTue 15:00 — 1 commitsTue 16:00 — 0 commitsTue 17:00 — 0 commitsTue 18:00 — 1 commitsTue 19:00 — 1 commitsTue 20:00 — 1 commitsTue 21:00 — 1 commitsTue 22:00 — 0 commitsTue 23:00 — 2 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 1 commitsWed 8:00 — 1 commitsWed 9:00 — 2 commitsWed 10:00 — 11 commitsWed 11:00 — 5 commitsWed 12:00 — 11 commitsWed 13:00 — 3 commitsWed 14:00 — 8 commitsWed 15:00 — 2 commitsWed 16:00 — 1 commitsWed 17:00 — 2 commitsWed 18:00 — 10 commitsWed 19:00 — 3 commitsWed 20:00 — 4 commitsWed 21:00 — 8 commitsWed 22:00 — 4 commitsWed 23:00 — 1 commitsThu 0:00 — 1 commitsThu 1:00 — 0 commitsThu 2:00 — 0 commitsThu 3:00 — 1 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 1 commitsThu 9:00 — 1 commitsThu 10:00 — 0 commitsThu 11:00 — 2 commitsThu 12:00 — 3 commitsThu 13:00 — 2 commitsThu 14:00 — 1 commitsThu 15:00 — 0 commitsThu 16:00 — 6 commitsThu 17:00 — 1 commitsThu 18:00 — 2 commitsThu 19:00 — 1 commitsThu 20:00 — 0 commitsThu 21:00 — 0 commitsThu 22:00 — 1 commitsThu 23:00 — 0 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 3 commitsFri 10:00 — 1 commitsFri 11:00 — 1 commitsFri 12:00 — 1 commitsFri 13:00 — 0 commitsFri 14:00 — 2 commitsFri 15:00 — 0 commitsFri 16:00 — 5 commitsFri 17:00 — 0 commitsFri 18:00 — 2 commitsFri 19:00 — 8 commitsFri 20:00 — 2 commitsFri 21:00 — 0 commitsFri 22:00 — 0 commitsFri 23:00 — 0 commitsSat 0:00 — 4 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 4 commitsSat 12:00 — 7 commitsSat 13:00 — 0 commitsSat 14:00 — 2 commitsSat 15:00 — 2 commitsSat 16:00 — 0 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
May 19, 2026daily#23+32
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    456.7K stars · TypeScript

  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • anomalyco/opencode

    The open source coding agent.

    211.7K stars · TypeScript

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript