antoinezambelli/forgePublic

A Python framework for self-hosted LLM tool-calling and multi-step agentic workflows

AI summary: A Python reliability layer for self-hosted LLM tool-calling and workflow constraint management.

Stars
2.2K
Forks
171
Watchers
13
Open issues
4
Open PRs
0
Contributors
~8
Commits
75
Branches
1

PythonMITCreated Feb 16, 2026Last push 3d agoLatest release v0.8.3+4 stars this week+4 this month

Star history

since May 17, 2026
01K2KMay 2026Jun 2026Jul 2026Aug 2026
2.2K stars as of Aug 7, 2026, tracked back to May 17, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 2 commits2026-03-16: 1 commit2026-03-17: 1 commit2026-03-18: 0 commits2026-03-19: 3 commits2026-03-20: 0 commits2026-03-21: 1 commit2026-03-22: 1 commit2026-03-23: 5 commits2026-03-24: 1 commit2026-03-25: 2 commits2026-03-26: 1 commit2026-03-27: 2 commits2026-03-28: 1 commit2026-03-29: 1 commit2026-03-30: 0 commits2026-03-31: 3 commits2026-04-01: 2 commits2026-04-02: 4 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 1 commit2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 2 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 1 commit2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 1 commit2026-04-30: 0 commits2026-05-01: 1 commit2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 0 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 1 commit2026-05-23: 2 commits2026-05-24: 3 commits2026-05-25: 1 commit2026-05-26: 3 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 1 commit2026-05-30: 0 commits2026-05-31: 6 commits2026-06-01: 1 commit2026-06-02: 1 commit2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 1 commit2026-06-12: 0 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 1 commit2026-06-17: 0 commits2026-06-18: 1 commit2026-06-19: 0 commits2026-06-20: 4 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 0 commits2026-06-27: 1 commit2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 0 commits2026-07-03: 2 commits2026-07-04: 1 commit2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 0 commits2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 2 commits2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 0 commits2026-07-18: 1 commit2026-07-19: 0 commits2026-07-20: 0 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 0 commits2026-07-29: 1 commit2026-07-30: 0 commits2026-07-31: 0 commits2026-08-01: 0 commits2026-08-02: 1 commit2026-08-03: 3 commits2026-08-04: 0 commits2026-08-05: 0 commits2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
75 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

What forge does

Forge is a specialized Python framework designed to make tool-calling highly reliable for self-hosted LLMs. It acts as a reliability layer inside an agentic loop, applying rescue parsing, retry nudges, and response validation to ensure models execute tools correctly. Developers can optionally enforce workflow structures using prerequisites and required steps to constrain the model's behavior. By focusing strictly on tool execution rather than orchestration, it significantly boosts the success rate of smaller local models (like 8B parameters). It is domain-agnostic and built to enhance existing single-agent setups rather than replace multi-agent graphs.

AI engineers and developers building robust single-agent systems using local or self-hosted LLMs. It requires proficiency in Python and an understanding of function calling APIs.

  • Feature: Enforces reliability for LLM tool-calling through rescue parsing and retry nudges.
  • Feature: Allows developers to define strict workflow constraints like required steps and prerequisites.
  • Feature: Validates model responses to ensure they match expected schemas before execution.
  • Feature: Specifically optimized to improve the tool-calling success rates of smaller, self-hosted models.
  • Feature: Operates as a domain-agnostic layer that fits into existing single-agent loops.

Where teams use it

Local Agent Reliability

Dramatically improves the tool-calling accuracy of 8B parameter local models running on constrained hardware.

Strict Workflow Execution

Forces an LLM to complete a specific sequence of validation steps before returning a final answer.

Error Recovery

Automatically catches malformed JSON or hallucinated tool names and nudges the model to correct them.

Custom Agent Enhancement

Integrates into custom Python agents to handle the brittle parsing layer of LLM responses.

Getting started: pip install forge

README

main branch

forge

PyPI Tests codecov Python 3.12+ License: MIT

A reliability layer for self-hosted LLM tool-calling. You give forge a set of tools; the model calls whichever it wants in whatever order. Workflow structure is opt-in — required_steps, prerequisites, and terminal_tool let you constrain the loop when you need to, but forge's guardrails (rescue parsing, retry nudges, response validation) apply with zero required steps too.

Forge takes an 8B local model from single digits to 84% across forge's 26-scenario v0.7.0 eval suite — and even lifts Sonnet 4.6 from 85% to 98% on the same workload (Anthropic numbers measured in v0.6.0; not re-run in v0.7.0 since the cost is non-trivial).

What forge isn't:

  • Not an agent orchestrator. Forge sits inside one agentic loop and makes its tool calls reliable. Multi-agent graphs, DAG planners, and cross-agent coordination are out of scope.
  • Not a coding harness. Forge is domain-agnostic. If you're building a coding agent (or already using one like opencode, aider, Cline), proxy mode lifts your existing harness with forge's guardrails — no rewrite.

Three ways to use it:

  • Proxy server — Drop-in proxy (python -m forge.proxy) speaking both the OpenAI chat-completions and Anthropic Messages (/v1/messages) APIs, sitting between any client and a local model server. Point OpenAI-compatible tools (opencode, Continue, aider) or Claude Code at it and forge applies guardrails transparently — the client thinks it's talking to a smarter model. Most popular entry point.

  • WorkflowRunner — Define tools, pick a backend, run structured agent loops. Forge manages the full lifecycle: system prompts, tool execution, context compaction, and guardrails. SlotWorker adds priority-queued access to a shared inference slot with auto-preemption — for multi-agent architectures where specialist workflows share a GPU slot. Best when you're building on forge directly.

  • Guardrails middleware — Use forge's reliability stack (composable middleware) inside your own orchestration loop. You control the loop; forge validates responses, rescues malformed tool calls, and enforces required steps.

Supports Ollama, llama-server (llama.cpp), Llamafile, vLLM, and Anthropic as backends.

Requirements

  • Python 3.12+
  • A running LLM backend (see below)

Install

pip install forge-guardrails                # core only
pip install "forge-guardrails[anthropic]"   # + Anthropic client

For development:

git clone https://github.com/antoinezambelli/forge.git
cd forge
pip install -e ".[dev]"

Backend setup (pick one)

llama-server (recommended — top 10 eval configs all run on llama-server):

# Install from https://github.com/ggml-org/llama.cpp/releases
llama-server -m path/to/Ministral-3-8B-Instruct-2512-Q8_0.gguf --jinja -ngl 999 --port 8080

Ollama (alternative — easier setup, slightly weaker on harder workloads):

# Install from https://ollama.com/download
ollama pull ministral-3:8b-instruct-2512-q4_K_M

Anthropic (API, no local GPU needed):

pip install -e ".[anthropic]"
export ANTHROPIC_API_KEY=sk-...

See Backend Setup for full instructions and Model Guide for which model fits your hardware.

Quick Start

Start llama-server however you normally do (e.g. in a separate shell):

llama-server -m path/to/Ministral-3-8B-Instruct-2512-Q8_0.gguf --jinja -ngl 999 --port 8080

Then the Python you'll run (e.g. from another shell):

import asyncio
from pydantic import BaseModel, Field
from forge import (
    Workflow, ToolDef, ToolSpec,
    WorkflowRunner, LlamafileClient,
    ContextManager, TieredCompact,
)

def get_weather(city: str) -> str:
    return f"72°F and sunny in {city}"

class GetWeatherParams(BaseModel):
    city: str = Field(description="City name")

workflow = Workflow(
    name="weather",
    description="Look up weather for a city.",
    tools={
        "get_weather": ToolDef(
            spec=ToolSpec(
                name="get_weather",
                description="Get current weather",
                parameters=GetWeatherParams,
            ),
            callable=get_weather,
        ),
    },
    required_steps=[],
    terminal_tool="get_weather",
    system_prompt_template="You are a helpful assistant. Use the available tools to answer the user.",
)

async def main():
    client = LlamafileClient(
        gguf_path="path/to/Ministral-3-8B-Instruct-2512-Q8_0.gguf",
        mode="native",
        recommended_sampling=True,
    )
    ctx = ContextManager(strategy=TieredCompact(keep_recent=2), budget_tokens=8192)
    runner = WorkflowRunner(client=client, context_manager=ctx)
    await runner.run(workflow, "What's the weather in Paris?")

asyncio.run(main())

For multi-step workflows, multi-turn conversations, and backend auto-management, see the User Guide. If you're building a long-running session (CLI, chat server, voice assistant), see the long-running session advisory for important guidance on filtering transient messages.

Proxy Server

Drop-in proxy that sits between any client and a local model server, speaking both the OpenAI chat-completions API and the Anthropic Messages API (/v1/messages). Point your client at the proxy (e.g. http://localhost:8081/v1) and forge applies its guardrails transparently — the client thinks it's talking to a smarter model.

This is the path for using forge with an existing harness (opencode, Continue, aider, Cline, anything that speaks the OpenAI chat-completions schema — or Claude Code, which speaks the Anthropic Messages API). No Python rewrite. Reasoning replay defaults to none: Forge still captures reasoning for observability, but keeps it out of backend-facing history on later turns — the most token-efficient policy, and statistically indistinguishable from replay-all on the eval suite (see reasoning-replay results). Use --reasoning-replay keep-last to replay only the latest reasoning block, or --reasoning-replay full for the historical replay-all behavior.

# External mode — you manage the backend, forge proxies it
python -m forge.proxy --backend-url http://localhost:8080 --port 8081

# Managed mode — forge starts the backend and the proxy together
python -m forge.proxy --backend llamaserver --gguf path/to/model.gguf --port 8081

# Managed vLLM — pass a model directory or HF repo id via --model-path
python -m forge.proxy --backend vllm --model-path /path/to/awq-dir --port 8081

Then configure your client to use http://localhost:8081/v1 as the API base URL.

Claude Code: the proxy also serves the Anthropic Messages API on POST /v1/messages, so you can point Claude Code at a forge-guarded local model — set ANTHROPIC_BASE_URL=http://localhost:8081 and ANTHROPIC_AUTH_TOKEN=anything for the claude process. See Using forge with Claude Code for the full setup (native-vs-prompt FC, Anthropic-shape downstreams, cache_control).

Backend compatibility:

  • Managed mode spins up the backend for you. Supported backends: llamaserver, llamafile, ollama, vllm (use --backend <name> with --gguf for the GGUF-based backends, --model-path for vllm, or --model for ollama).
  • External mode is backend-agnostic — forge talks POST /v1/chat/completions to whatever you point --backend-url at, as long as it speaks the OpenAI schema. Tool calls must come back in OpenAI tool_calls format or in one of forge's rescue-parsed formats (Mistral [TOOL_CALLS], Qwen <tool_call> XML, fenced JSON). For a vLLM server, add --backend vllm so the proxy adopts vLLM's --served-model-name (vLLM 404s on a mismatched model field, unlike llama.cpp). An explicit --model overrides that discovery — for hosted multi-model gateways (one /v1/models endpoint listing many models), pin your model and skip discovery entirely: --backend vllm --model <name> --budget-tokens <n> --backend-api-key <key>.

What proxy mode fortifies

On every POST /v1/chat/completions, forge applies (in order):

  1. Response validation — each tool call in the model's response is checked against the tools array in the request. Calls to unknown tool names or with malformed shapes are caught before the response returns to your client.
  2. Rescue parsing — when the model emits tool calls in the wrong format (JSON in a code fence, Mistral's [TOOL_CALLS]name{args}, Qwen's <tool_call>...</tool_call> XML), forge extracts the structured call and re-emits it in the canonical OpenAI tool_calls schema. Biggest practical lift for Mistral-family models.
  3. Retry loop with error tracking — if validation fails, forge retries inference up to --max-retries (default 3) with a corrective tool-result message on the canonical channel, rather than returning a malformed response. From your client's perspective the proxy looks like a single request that just took a few extra ms.
  4. Synthetic respond tool injection — when tools are present in the request, forge injects a synthetic respond tool the model calls instead of producing bare text. The respond call is stripped from the outbound response — the client sees a normal text response (finish_reason: "stop") and never knows the tool exists. Essential for small local models (~8B) that can't be trusted to choose correctly between text and tool calls. See ADR-013 for the full analysis.

What proxy mode does not do

Proxy mode is single-shot per request; some forge features need multi-turn workflow state that the OpenAI chat-completions schema doesn't carry:

  • Prerequisite enforcement and step-ordering — these need a workflow definition spanning turns. Available in WorkflowRunner.
  • Context compaction and session memory — proxy mode forwards the inbound message list as-is; managing the rolling window is the client's job.
  • VRAM-aware budget detection — opt in with --budget-mode forge-full or --budget-mode forge-fast; otherwise proxy uses the backend's reported budget.

For the full guardrail surface, use WorkflowRunner directly. The proxy trades depth for "use forge with your existing setup, no rewrite."

Useful flags

Flag Default Purpose
--max-retries N 3 Retry budget per validation failure
--no-rescue (rescue on) Disable rescue parsing (debugging only)
--budget-mode {backend,manual,forge-full,forge-fast} backend Context budget source
--budget-tokens N Manual token budget (requires --budget-mode manual)
--serialize / --no-serialize auto Force request serialization (single-slot backends)

Docker

You can run the forge proxy as a Docker container.

Build the image:

docker build -t forge-proxy .

Run the container:

# Connect to an external backend (e.g. vLLM hosted on the same machine)
docker run -p 8081:8081 forge-proxy --backend-url http://host.docker.internal:8000 --backend vllm --budget-mode manual --budget-tokens 8192

Note: If your backend is running on localhost of the host machine, use http://host.docker.internal:PORT (on macOS/Windows) or the host's IP address to allow the container to reach it.

Backends

Backend Best for Native FC?
Ollama Easiest setup, model management built-in Yes
llama-server Best performance, full control Yes (with --jinja)
Llamafile Single binary, zero dependencies No (prompt-injected)
vLLM High-throughput serving, AWQ/GPTQ weights Yes (server-side parser)
Anthropic Frontier baseline, hybrid workflows Yes

See Backend Setup for installation and Model Guide for which model to pick.

Running Tests

python -m pytest tests/ -v --tb=short
python -m pytest tests/ --cov=forge --cov-report=term-missing

Eval Harness

26 scenarios measuring how reliably a model + backend combo navigates multi-step tool-calling workflows — split into an OG-18 baseline tier and an 8-scenario advanced_reasoning tier for top-end separation. See Eval Guide for full CLI reference.

# llama-server (start in another terminal first; see Eval Guide)
python -m tests.eval.eval_runner --backend llamafile --llamafile-mode prompt --gguf "path/to/Ministral-3-8B-Instruct-2512-Q8_0.gguf" --runs 10 --stream --verbose

# Batch eval (JSONL output, automatic resume)
python -m tests.eval.batch_eval --config all --runs 50

# Reports — ASCII table by default; --html / --markdown export views
python -m tests.eval.report eval_results.jsonl
python -m tests.eval.report eval_results.jsonl --html docs/results/dashboard.html
python -m tests.eval.report eval_results.jsonl --markdown docs/results/

Project Structure

src/forge/
  __init__.py          # Public API exports
  errors.py            # ForgeError hierarchy
  server.py            # setup_backend(), ServerManager, BudgetMode
  core/
    messages.py        # Message, MessageRole, MessageType, MessageMeta
    workflow.py        # ToolSpec, ToolDef, ToolCall, TextResponse, Workflow
    inference.py       # run_inference() — shared front half (compact, fold, validate, retry)
    runner.py          # WorkflowRunner — the agentic loop
    slot_worker.py     # SlotWorker — priority-queued slot access
    steps.py           # StepTracker
  guardrails/
    guardrails.py      # Guardrails facade — applies the full stack in foreign loops
    nudge.py           # Nudge dataclass
    response_validator.py  # ResponseValidator, ValidationResult
    step_enforcer.py   # StepEnforcer, StepCheck
    error_tracker.py   # ErrorTracker
  clients/
    base.py            # ChunkType, StreamChunk, LLMClient protocol
    ollama.py          # OllamaClient (native FC)
    llamafile.py       # LlamafileClient (native FC or prompt-injected)
    anthropic.py       # AnthropicClient (frontier baseline)
  context/
    manager.py         # ContextManager, CompactEvent
    strategies.py      # CompactStrategy, NoCompact, TieredCompact, SlidingWindowCompact
    hardware.py        # HardwareProfile, detect_hardware()
  prompts/
    templates.py       # Tool prompt builders (prompt-injected path)
    nudges.py          # Retry and step-enforcement nudge templates
  tools/
    respond.py         # Synthetic respond tool (respond_tool(), respond_spec())
  proxy/
    __main__.py        # CLI entry point: python -m forge.proxy
    proxy.py           # ProxyServer — programmatic start/stop API
    server.py          # Raw asyncio HTTP server, SSE streaming
    handler.py         # Request handler — bridge between HTTP and run_inference
    convert.py         # OpenAI messages ↔ forge Messages conversion
tests/
  unit/                # 865 deterministic tests — no LLM backend required
  eval/                # Eval harness — model qualification against real backends

Documentation

  • User Guide — Usage patterns, multi-turn, context management, guardrails, slot worker, long-running session advisory
  • Model Guide — Which model and backend for your hardware
  • Backend Setup — Backend installation and server setup
  • Eval Guide — Eval harness CLI reference, batch eval
  • Architecture — Full design document
  • Workflow Internals — Workflow design and runner internals
  • Contributing — How to set up, test, and add new backends or scenarios

Paper

The forge guardrail framework and ablation study are published as:

Zambelli, A. Forge: A Reliability Layer for Self-Hosted LLM Tool-Calling. https://doi.org/10.1145/3786335.3813193

A pre-publication preprint is also available at docs/forge_ieee_preprint.pdf — kept as a historical artifact. Cite the published version above; the DOI link may not resolve immediately depending on the publisher's release timing.

License

MIT — Copyright (c) 2025-2026 Antoine Zambelli

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

11 total
  1. ## [0.8.3] — 2026-08-03 A proxy-correctness and evaluation-publication release. Forge restores model identity handling across vLLM and Anthropic proxy paths while making evaluation collection and publication deterministic and provenance-safe. ### Added - **Deterministic Hugging Face dataset publication tooling.** A new offline `tests.eval.dataset_builder` command produces verified Parquet bundles with `latest`, `snapshot`, and `history` views, a fixed schema, source provenance, content hashes, and an upload-ready Dataset Card. Builds are verified and published atomically, with PyArrow available through the dedicated `dataset-builder` extra. #136 ### Changed - **Eval-generation provenance is explicit and resume-safe.** `batch_eval --generation` stamps every collected row with its comparability epoch. Existing JSONL outputs are validated before collection begins, rejecting malformed data, mixed or mismatched generations, and ambiguous resume identities. Shared generation and replay-policy semantics now govern collection, reporting, and publication. #135 ### Fixed - **Unpinned external vLLM proxies discover their served model before listing it.** The first `GET /v1/models` can use

  2. ## [0.8.2] — 2026-07-28 A maintenance and eval-publication release. Forge hardens malformed llama.cpp 500 recovery, adds compatibility with clients that use llama-server’s unversioned chat endpoint, and publishes an expanded evaluation dashboard covering 378,300 runs across 291 configurations. ### Added - **`POST /chat/completions` proxy alias.** The proxy now serves llama-server’s unversioned endpoint as an alias of `/v1/chat/completions`, restoring compatibility with clients such as pi-llama-cpp. - **Expanded large-model eval tiers.** The published roster gains Gemma-4 large models, new Qwen3.6 replay sweeps, LFM2.5 and Mellum2 replay ladders, and a new 120B tier. The consolidated v0.8.2 dataset contains 74,100 generation-3 runs across 57 complete arms. - **Reasoning effort as an eval dimension.** Batch resume identity, result rows, report deduplication, and display labels now distinguish effort variants of the same model without conflating their results. ### Changed - **Evaluation dashboard regenerated across all published datasets.** The dashboard now represents 378,300 runs and 291 model/backend configurations, with normalized family names for the new models. - **Replay-awa

  3. A bug-fix release for the llamafile backend. When llama.cpp's tool-call parser rejects malformed model output with a 500, the raw error JSON no longer leaks into the conversation as assistant text — complete tool calls are rescued out of the error body and executed, and unrecoverable ones trigger a clean re-sample nudge. ### Added - **Tool-call rescue from malformed-500 bodies.** llama.cpp's `Failed to parse input` message embeds the rejected generation; forge now leniently re-parses `<tool_call>` blocks out of it (Qwen-coder XML format) and returns each block that names a tool from the request's `tools` array as a real `ToolCall` — deduped, with parameters coerced to their declared schema types. Skeleton/preview blocks are passed through too: dispatch rejects them with a `[ToolError]` on the tool channel, the canonical corrective signal, while complete calls simply execute. Unknown tool names are never fabricated. Rescues are counted on `LlamafileClient.rescued_tool_calls` and logged, so rescued runs stay auditable. ### Fixed - **Malformed tool-call 500s no longer leak error JSON into the conversation.** When llama.cpp rejects a malformed or incomplete tool call and nothing can

  4. ## [0.8.0] — 2026-06-27 First-class authentication across all proxy modes and backends. forge now forwards exactly one credential to the backend — a static `--backend-api-key` or a single inbound auth header — relocating it across protocols (`x-api-key` ↔ `Authorization: Bearer`) when the frontend and backend differ. Gated OpenAI-compatible backends (LM Studio, hosted vLLM, service accounts) work without monkey-patching. Closes #119. ### Added - **`--backend-api-key` / `FORGE_BACKEND_API_KEY`** — a static credential forge sends to the backend in its native auth slot (LM Studio, hosted providers, service accounts — the case where the caller sends nothing). Baked into the backend client at startup and relocated to the backend's protocol slot. When set, an inbound auth header is refused as a second credential. - **Cross-protocol credential relocation** — an inbound `x-api-key` ↔ `Authorization: Bearer` is rewritten to the backend's protocol when frontend and backend differ (the SSO/forwarded-token case). Frontend protocol is by path (`/v1/chat/completions` = openai, `/v1/messages` = anthropic); backend by `--backend-protocol`. See the auth section in [Backend Setup](docs/BACKEND_SET

  5. ## [0.7.6] — 2026-06-20 A bug-fix release for the Ollama backend and inline reasoning capture. Multi-turn tool sessions and multi-part message content no longer 400 against Ollama's native API, and chain-of-thought emitted inline in `content` is now captured on vLLM and Ollama as it already was on the structured-field path. ### Added - **16GB-tier MoE models** in the published eval set and dashboard (gen-3 regeneration). #107 ### Changed - **Think-tag parsing consolidated** into one shared helper (`forge.prompts.think_tags`), de-duplicating inline-reasoning extraction across the llamafile client and prompt templates. #112 - **Scripted test doubles consolidated** into a shared `conftest` fixture, replacing the per-module `MockClient` stand-ins. #76 (thanks @SuperMarioYL). ### Fixed - **Inline `<think>` reasoning is captured on vLLM and Ollama.** When a reasoning model emits its chain-of-thought inline in `content` (`<think>…</think>`) instead of a structured reasoning field, that reasoning is now extracted onto the first tool call — matching the behavior already present for structured reasoning fields. #110 - **Ollama's native `/api/chat` accepts OpenAI-wire message shapes.** On

Code frequency

additions and deletions
+56.4K-56.4KWeek of 2026-03-15: +56,362 linesWeek of 2026-03-15: -2,100 linesWeek of 2026-03-22: +2,556 linesWeek of 2026-03-22: -353 linesWeek of 2026-03-29: +2,237 linesWeek of 2026-03-29: -23,030 linesWeek of 2026-04-05: +1,452 linesWeek of 2026-04-05: -381 linesWeek of 2026-04-12: +143 linesWeek of 2026-04-12: -1 linesWeek of 2026-04-19: +2,742 linesWeek of 2026-04-19: -2,350 linesWeek of 2026-04-26: +8,261 linesWeek of 2026-04-26: -2,590 linesWeek of 2026-05-03: +0 linesWeek of 2026-05-03: -0 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +1,567 linesWeek of 2026-05-17: -2,907 linesWeek of 2026-05-24: +4,867 linesWeek of 2026-05-24: -581 linesWeek of 2026-05-31: +3,400 linesWeek of 2026-05-31: -705 linesWeek of 2026-06-07: +2,762 linesWeek of 2026-06-07: -765 linesWeek of 2026-06-14: +863 linesWeek of 2026-06-14: -195 linesWeek of 2026-06-21: +3,187 linesWeek of 2026-06-21: -129 linesWeek of 2026-06-28: +1,041 linesWeek of 2026-06-28: -52 linesWeek of 2026-07-05: +840 linesWeek of 2026-07-05: -60 linesWeek of 2026-07-12: +21 linesWeek of 2026-07-12: -1 linesWeek of 2026-07-19: +0 linesWeek of 2026-07-19: -0 linesWeek of 2026-07-26: +0 linesWeek of 2026-07-26: -0 linesMar 15, 2026Jul 26, 2026
+92.3K lines added, -36.2K removed over the last year.

Commits per week

last 52 weeks
130Week of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 8 commitsWeek of 2026-03-22: 13 commitsWeek of 2026-03-29: 10 commitsWeek of 2026-04-05: 1 commitsWeek of 2026-04-12: 2 commitsWeek of 2026-04-19: 1 commitsWeek of 2026-04-26: 2 commitsWeek of 2026-05-03: 0 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 3 commitsWeek of 2026-05-24: 8 commitsWeek of 2026-05-31: 8 commitsWeek of 2026-06-07: 1 commitsWeek of 2026-06-14: 6 commitsWeek of 2026-06-21: 1 commitsWeek of 2026-06-28: 3 commitsWeek of 2026-07-05: 2 commitsWeek of 2026-07-12: 1 commitsWeek of 2026-07-19: 0 commitsWeek of 2026-07-26: 1 commitsWeek of 2026-08-02: 4 commitsAug 10, 2025Aug 2, 2026
75 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 1 commitsSun 3:00 — 0 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 1 commitsSun 11:00 — 0 commitsSun 12:00 — 0 commitsSun 13:00 — 2 commitsSun 14:00 — 2 commitsSun 15:00 — 1 commitsSun 16:00 — 0 commitsSun 17:00 — 0 commitsSun 18:00 — 0 commitsSun 19:00 — 1 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 3 commitsSun 23:00 — 3 commitsMon 0:00 — 1 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 0 commitsMon 8:00 — 1 commitsMon 9:00 — 0 commitsMon 10:00 — 0 commitsMon 11:00 — 0 commitsMon 12:00 — 1 commitsMon 13:00 — 0 commitsMon 14:00 — 0 commitsMon 15:00 — 0 commitsMon 16:00 — 1 commitsMon 17:00 — 2 commitsMon 18:00 — 1 commitsMon 19:00 — 0 commitsMon 20:00 — 1 commitsMon 21:00 — 0 commitsMon 22:00 — 2 commitsMon 23:00 — 1 commitsTue 0:00 — 0 commitsTue 1:00 — 1 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 3 commitsTue 6:00 — 1 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 0 commitsTue 12:00 — 0 commitsTue 13:00 — 0 commitsTue 14:00 — 0 commitsTue 15:00 — 0 commitsTue 16:00 — 1 commitsTue 17:00 — 2 commitsTue 18:00 — 0 commitsTue 19:00 — 0 commitsTue 20:00 — 0 commitsTue 21:00 — 1 commitsTue 22:00 — 1 commitsTue 23:00 — 1 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 1 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 0 commitsWed 9:00 — 0 commitsWed 10:00 — 0 commitsWed 11:00 — 0 commitsWed 12:00 — 0 commitsWed 13:00 — 0 commitsWed 14:00 — 0 commitsWed 15:00 — 0 commitsWed 16:00 — 0 commitsWed 17:00 — 0 commitsWed 18:00 — 0 commitsWed 19:00 — 2 commitsWed 20:00 — 0 commitsWed 21:00 — 0 commitsWed 22:00 — 1 commitsWed 23:00 — 1 commitsThu 0:00 — 1 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 1 commitsThu 4:00 — 1 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 0 commitsThu 11:00 — 0 commitsThu 12:00 — 0 commitsThu 13:00 — 0 commitsThu 14:00 — 0 commitsThu 15:00 — 0 commitsThu 16:00 — 1 commitsThu 17:00 — 0 commitsThu 18:00 — 1 commitsThu 19:00 — 0 commitsThu 20:00 — 2 commitsThu 21:00 — 0 commitsThu 22:00 — 1 commitsThu 23:00 — 1 commitsFri 0:00 — 1 commitsFri 1:00 — 1 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 2 commitsFri 8:00 — 0 commitsFri 9:00 — 1 commitsFri 10:00 — 1 commitsFri 11:00 — 0 commitsFri 12:00 — 0 commitsFri 13:00 — 0 commitsFri 14:00 — 0 commitsFri 15:00 — 0 commitsFri 16:00 — 0 commitsFri 17:00 — 0 commitsFri 18:00 — 3 commitsFri 19:00 — 1 commitsFri 20:00 — 0 commitsFri 21:00 — 0 commitsFri 22:00 — 1 commitsFri 23:00 — 0 commitsSat 0:00 — 1 commitsSat 1:00 — 0 commitsSat 2:00 — 2 commitsSat 3:00 — 0 commitsSat 4:00 — 1 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 1 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 1 commitsSat 14:00 — 0 commitsSat 15:00 — 0 commitsSat 16:00 — 0 commitsSat 17:00 — 3 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 2 commitsSat 22:00 — 0 commitsSat 23:00 — 1 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits64 (85%)
Community commits11 (15%)

75 commits in total over the last year.

DateListRankStars gained
May 20, 2026daily#12+58
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python

  • awesome-selfhosted/awesome-selfhosted

    A list of Free Software network services and web applications which can be hosted on your own servers

    311.2K stars

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    277.2K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    238.5K stars · JavaScript

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    234.7K stars · JavaScript