tashfeenahmed/freellmapiPublic

OpenAI-compatible proxy that stacks the free tiers of 28 LLM providers (~4B tokens/month) behind one /v1 endpoint — plus any custom OpenAI-compatible endpoint. Smart routing, automatic failover, encrypted keys. Personal experimentation only.

AI summary: A unified OpenAI-compatible router that aggregates free-tier quotas from dozens of LLM providers into a single API endpoint.

Stars
17.8K
+133 today
Forks
2.6K
Watchers
73
Open issues
18
Open PRs
31
Contributors
~54
Commits
502
Branches
3

TypeScriptMITCreated Apr 21, 2026Last push 1d agoLatest release v0.6.8+544 stars this week+544 this month

Star history

since Apr 19, 2026
05K10K15KApr 2026May 2026Jun 2026Aug 2026
17.8K stars as of Aug 6, 2026, tracked back to Apr 19, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 1 commit2026-04-22: 6 commits2026-04-23: 1 commit2026-04-24: 0 commits2026-04-25: 1 commit2026-04-26: 1 commit2026-04-27: 0 commits2026-04-28: 2 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 4 commits2026-05-02: 7 commits2026-05-03: 1 commit2026-05-04: 2 commits2026-05-05: 0 commits2026-05-06: 0 commits2026-05-07: 0 commits2026-05-08: 0 commits2026-05-09: 0 commits2026-05-10: 2 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 1 commit2026-05-18: 0 commits2026-05-19: 3 commits2026-05-20: 3 commits2026-05-21: 0 commits2026-05-22: 4 commits2026-05-23: 9 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 9 commits2026-05-27: 2 commits2026-05-28: 2 commits2026-05-29: 1 commit2026-05-30: 4 commits2026-05-31: 14 commits2026-06-01: 2 commits2026-06-02: 12 commits2026-06-03: 4 commits2026-06-04: 18 commits2026-06-05: 26 commits2026-06-06: 3 commits2026-06-07: 8 commits2026-06-08: 7 commits2026-06-09: 2 commits2026-06-10: 10 commits2026-06-11: 7 commits2026-06-12: 12 commits2026-06-13: 6 commits2026-06-14: 3 commits2026-06-15: 2 commits2026-06-16: 1 commit2026-06-17: 3 commits2026-06-18: 0 commits2026-06-19: 4 commits2026-06-20: 5 commits2026-06-21: 8 commits2026-06-22: 5 commits2026-06-23: 0 commits2026-06-24: 2 commits2026-06-25: 0 commits2026-06-26: 2 commits2026-06-27: 11 commits2026-06-28: 6 commits2026-06-29: 0 commits2026-06-30: 5 commits2026-07-01: 0 commits2026-07-02: 14 commits2026-07-03: 3 commits2026-07-04: 0 commits2026-07-05: 11 commits2026-07-06: 10 commits2026-07-07: 11 commits2026-07-08: 0 commits2026-07-09: 2 commits2026-07-10: 0 commits2026-07-11: 13 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 3 commits2026-07-15: 2 commits2026-07-16: 0 commits2026-07-17: 3 commits2026-07-18: 0 commits2026-07-19: 5 commits2026-07-20: 4 commits2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 2 commits2026-07-26: 14 commits2026-07-27: 9 commits2026-07-28: 18 commits2026-07-29: 10 commits2026-07-30: 6 commits2026-07-31: 6 commits2026-08-01: 5 commits2026-08-02: 9 commits2026-08-03: 7 commits2026-08-04: 0 commits2026-08-05: 13 commits2026-08-06: 3 commits2026-08-07: 0 commits2026-08-08: 0 commits
437 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    17,831 stars

  • Breakout launch

    17,831 stars in 108 days

  • Actively maintained

    Pushed within 48 hours

  • Well documented

    High community health score

  • Permissive license

    MIT

  • Continuous integration

    Automated checks passing

  • Repeat trending

    3 trending appearances

What freellmapi does

FreeLLMAPI acts as a sophisticated proxy and router, combining the free tiers of 29 different LLM providers into one accessible `/v1` endpoint. It solves the problem of managing multipleAPI keys and hitting rate limits by automatically falling over to the next available provider when a quota is exhausted. The system encrypts all stored keys and meticulously tracks per-key usage to ensure users never exceed their free-tier caps. It maintains an automated model catalog via a signed feed, keeping track of over 350 free endpoints across various model families. By presenting a standard OpenAI-compatible interface, it allows developers to integrate massive amounts of free inference compute into their applications without changing their client code.

Developers, researchers, and hobbyists looking to maximize free LLM usage and build resilient, multi-provider AI applications. Requires basic knowledge of API integration and API key management.

  • OpenAI API compatibility: Exposes a standard `/v1` interface that works natively with existing OpenAI SDKs and clients.
  • Intelligent rate limit failover: Automatically reroutes requests to alternative providers the moment a specific free tier limit is hit.
  • Usage quota tracking: Monitors token consumption on a per-key basis to strictly enforce free tier boundaries and prevent unexpected charges.
  • Encrypted credential storage: Secures all provider API keys using robust encryption before storing them in the router configuration.
  • Dynamic model cataloging: Updates its internal registry of available models and endpoints automatically via a signed data feed.

Where teams use it

Zero-cost inference scaling

Indie developers can pool multiple free tier accounts to run high-volume LLM applications without incurring inference costs.

Resilient API routing

Applications requiring high uptime can use the router to automatically failover when their primary LLM provider experiences downtime.

Unified model testing

Researchers can query hundreds of different model families through a single API endpoint to evaluate performance differences.

Simplified key management

Teams can centralize their various provider keys into one secure proxy, distributing a single endpoint to their internal applications.

Getting started: Visit https://freellmapi.co/

README

main branch

FreeLLMAPI

4 billion tokens per month. 29 free LLM providers. 358 free model endpoints. One OpenAI-compatible endpoint.

Aggregate free tiers from dozens of providers, plus custom OpenAI-compatible chat, embedding, image, and audio endpoints, behind a single /v1 API. Keys are stored encrypted. A router picks the best available model for each request, falls over to the next provider when one is rate-limited, and tracks per-key usage so you stay under every free-tier cap.

CI GitHub stars License: MIT PRs Welcome Docker image Ask DeepWiki

freellmapi.co · browse the full catalog: 251 model families, 358 free endpoints

English · 简体中文

FreeLLMAPI dashboard — Models page with the monthly token budget

Your router updates its own model catalog from a signed feed: new free models, quota changes, and compatibility fixes land without a git pull. Go live at freellmapi.co ($19/yr, cancel anytime).


Contents

Guides: Install & deploy · API reference · Clients & coding agents · Prompt compression · Architecture & internals · Documentation index · Contributor guide

Why this exists

Every serious AI lab now offers a free tier, a few million tokens a month, a few thousand requests a day. On its own each tier is a toy. Stacked together, they add up to roughly 4 billion tokens per month of working inference capacity, across 251 model families / 358 provider endpoints from small-and-fast to reasonably capable.

The problem is that stacking them by hand is painful: twenty-nine different SDKs, twenty-nine different rate limits, twenty-nine places a request can fail. FreeLLMAPI collapses that into one OpenAI-compatible endpoint. Point any OpenAI client library at your local server, and it routes transparently across whichever providers you've added keys for.

And the free-tier landscape shifts weekly: providers launch models, retire them, and change quotas without notice. FreeLLMAPI tracks all of that for you. The router pulls a signed model catalog from freellmapi.co on its own, so your install keeps up without a git pull. See Premium (live catalog) for how fast it keeps up.

The free tier, stacked — ~4B tokens of free inference per month across 28 providers

Supported providers

Google
Google
Groq
Groq
Cerebras
Cerebras
OpenCode Zen
OpenCode Zen
Mistral
Mistral
OpenRouter
OpenRouter
Cloudflare
Cloudflare
Cohere
Cohere
Z.ai (Zhipu)
Z.ai (Zhipu)
NVIDIA
NVIDIA
HuggingFace
HuggingFace
ModelScope
Qwen3 · DeepSeek V4 · GLM-5 (needs Aliyun cn binding)

… and 17 more free providers

Plus a custom provider — point chat, embedding, image, or audio models at any OpenAI-compatible endpoint (llama.cpp, LM Studio, vLLM, a local Ollama, or a remote gateway) from the Keys page.

The full, always-current list lives at freellmapi.co/models with per-model rate limits, context windows, and free-token budgets.

Compatible CLIs & coding agents

Claude Code
Claude Code
Codex CLI
Codex CLI
Gemini CLI
Gemini CLI
Aider
Aider
Cline
Cline
Roo Code
Roo Code
Continue
Continue
OpenCode
OpenCode
Goose
Goose
Qwen Code
Qwen Code
Kilo Code
Kilo Code
Crush
Crush
Cursor
Cursor
Zed
Zed
JetBrains AI
JetBrains AI

… plus any OpenAI-compatible client, Anthropic SDK, Gemini SDK, or Ollama-capable app

Most of these configure themselves with one command — npx freellmapi setup-claude, setup-codex, setup-aider, and ten more generators that fetch your live catalog, back up existing config, and never clobber what's already there. Claude Code and Codex also get zero-persistence launchers (freellmapi launch, freellmapi launch-codex) that inject credentials into the child process only. Zed and JetBrains AI connect through the opt-in Ollama emulation; Gemini CLI speaks its native wire on /v1beta.

Per-tool recipes, the setup CLI reference, revocable URL tokens for headerless clients, and the MCP server all live in Clients & coding agents →

How it compares

Feature comparison against OpenRouter, LiteLLM, and Portkey

Based on public documentation, July 2026 — corrections welcome.

Features

Feature overview

  • Every OpenAI surface/v1/chat/completions, /v1/responses (what Codex CLI needs), /v1/completions (editor ghost-text autocomplete), /v1/images/generations, /v1/audio/speech, /v1/embeddings, and /v1/models — streaming and non-streaming, from the official SDKs or any OpenAI-compatible client. API reference →
  • Anthropic Messages API/v1/messages speaks Anthropic's wire format over the same router, so Claude Code and the official Anthropic SDKs run against your free pool. Details →
  • Native Gemini + Ollama surfaces — Gemini CLI can use /v1beta (generateContent, streaming, token counting, models), while opt-in Ollama emulation serves NDJSON chat/generate, tags, metadata, and embeddings for Zed, JetBrains, and other local-model clients.
  • Fusion (multi-model synthesis) — request the virtual fusion model and the router fans your prompt out to a panel of diverse free models in parallel, then a judge model synthesizes one answer from the drafts. Details →
  • Image generation & text-to-speech/v1/images/generations and /v1/audio/speech route across the providers that serve media models, including custom OpenAI-compatible media endpoints.
  • Tool calling & structured outputs — OpenAI-style tools round-trip across providers (plain-text tool calls are rescued into real tool_calls), plus response_format, seed, logprobs, penalties, and the rest of the sampling params passed through per provider.
  • Smart routing, six strategies — live per-model speed/capability/reliability scores rank your chain; automatic fallover retries the next model on 429/5xx with cooldowns and key rotation. Routing in detail →
  • Unified models & profiles — the same model on several providers collapses into one entry with strict in-group failover; named fallback-chain profiles (a coding chain, a vision chain) switch from the dashboard or per request via auto:<profile>.
  • Per-key rate tracking — RPM/RPD/TPM/TPD counters per (platform, model, key) that learn providers' reported ceilings, so routing always stays under every cap.
  • Self-updating model catalog — the router syncs a signed catalog from freellmapi.co twice a day: new models, quota changes, and provider quirk fixes land automatically. Premium →
  • Sticky sessions & context handoff — conversations stay on one model for 30 minutes; an optional compact handoff note keeps the thread coherent when a mid-chat switch does happen. Details →
  • Prompt compression (opt-in) — a shared, fail-open request pipeline can deduplicate prompts, filter tool output, compact repeated JSON, and trim stale context before cache lookup and routing. Details →
  • Encrypted keys, one token out — provider keys are AES-256-GCM encrypted in SQLite and decrypted in-memory per request; your apps only ever see a single unified freellmapi-… bearer token.
  • Admin dashboard & analytics — React UI to manage keys, reorder the chain, run a playground, and read p50/p95/TTFT analytics over 24h–90d windows; login-gated, dark/light themes, 60 languages.
  • MCP server & interactive docs — agents can introspect usable models, provider health, and routing strategy over /mcp; a dependency-free OpenAPI viewer lives at /v1/docs. Coding agents →
  • Ops niceties — opt-in response cache, encrypted DB backups, periodic key health checks, bulk key import/export, declarative startup config. Install & deploy →
  • Runs anywhere Node 20+ runs — Windows, macOS, Linux servers, or a small ARM SBC (Raspberry Pi included). ~40 MB RSS at idle behind PM2 / systemd / whatever supervisor you prefer.

The scope is deliberately narrow — see what's not supported yet.

Quick start

One-liner (Docker required — sets up ~/freellmapi, generates an encryption key, pulls the image, and starts the container):

curl -fsSL https://freellmapi.co/install.sh | bash

Prefer to read before you pipe to bash? The script is here. Re-running it is safe: your .env (and encryption key) is preserved and the container updates to :latest.

Open http://localhost:3001, add your provider keys on the Keys page, reorder the Fallback Chain to taste, and grab your unified API key from the Keys page header. That unified key is what you point your OpenAI SDK at.

On Windows, the easiest path is the desktop .exe installer from Releases (below). On Android, see the experimental Termux guide.

Everything else — Docker Compose, local development, declarative startup config, production builds, LAN access, and backups — is in docs/install.md.

Desktop app

A native menu-bar app lives in desktop/: the entire router + dashboard running locally from your tray, with a glass popover showing live request stats.

FreeLLMAPI desktop app

Download from Releases — the macOS .dmg and the Windows .exe installer are attached to every release. No account or password to set up: the only credential you need is the unified API key from the tray popover. Build-from-source steps and where your data lives: docs/install.md.

Works with OpenAI-compatible clients

Anything that can target an OpenAI-compatible base URL works: set it to http://localhost:3001/v1 with the unified key from the dashboard. Claude Code, Codex CLI, Cline / Roo Code, Continue (including inline autocomplete), Aider, opencode, and Cursor each have a short recipe in docs/clients.md — and the router doubles as an MCP server your agents can introspect mid-session.

The fastest setup is generated from the models available on your live server:

npx freellmapi setup-claude --url http://localhost:3001 --api-key <unified-key>

Every generator supports --dry-run, creates a timestamped backup before changing an existing file, and merges into the user's configuration. Launchers keep credentials out of config files entirely: npx freellmapi launch for Claude Code and npx freellmapi launch-codex for Codex.

Agent Automated setup Base URL
Claude Code setup-claude root
Codex CLI setup-codex /v1
Cline setup-cline /v1
Continue setup-continue /v1
Aider setup-aider /v1
OpenCode setup-opencode /v1
Goose setup-goose /v1
Qwen Code setup-qwen /v1 (or native /v1beta)
Roo / Kilo / Crush setup-roo / setup-kilo / setup-crush /v1
Cursor setup-cursor guide public /v1 URL

FreeLLMAPI is local-first and single-user by design. Your provider keys stay in your SQLite database, encrypted at rest, and requests go from your machine to the upstream providers you enabled.

Languages

The dashboard ships in 60 languages (the desktop tray menu in 6). The UI auto-detects your browser/system language on first load and you can switch any time from ⋯ → Settings; the choice is remembered. Right-to-left languages (العربية, עברית, فارسی, اردو) flip the whole layout automatically, and only the active language's dictionary is loaded — the rest never touch your bandwidth.

United States China Spain France Brazil Italy India Saudi Arabia Bangladesh Russia Pakistan Indonesia Germany Japan Kenya Türkiye Vietnam South Korea Iran Thailand Poland Ukraine Myanmar Romania Netherlands Malaysia Philippines Nigeria Ethiopia Uzbekistan Azerbaijan Sri Lanka Nepal Cambodia Greece Czechia Hungary Sweden Israel Denmark Finland Norway Slovakia Bulgaria Croatia Serbia Lithuania Taiwan Portugal Georgia

The full list of locales lives in client/src/i18n/locale-config.ts.

The original six locales are human-reviewed; the newer ones are machine- translated and improve as native speakers send corrections — a one-string PR is a great first contribution.

Translations live in client/src/i18n/locales/ as flat JSON files. To fix a string, edit the value in the locale's JSON file. To add a language, copy en.json, translate the values, and register the locale in client/src/i18n/locale-config.ts (and desktop/src/i18n.ts for the tray strings); npm test checks every locale for key/placeholder parity — PRs welcome.

Premium (live catalog)

The router keeps its model catalog fresh on its own: it pulls a signed catalog from freellmapi.co twice a day and applies new models, quota changes, and provider quirk fixes to your local DB. Your own enable/disable choices and custom providers are never touched, and every download is verified against a pinned Ed25519 key before it is applied.

The catalog currently tracks 29 providers, 251 model families, 358 provider/model endpoints, and roughly 4 billion tokens per month of listed free-tier capacity. Browse the full set at freellmapi.co/models.

Premium keeps that signed catalog live on every router you run. When a provider launches a strong free model, quietly tightens a quota, or breaks a wire format, live-feed routers receive the update as soon as we ship it.

Go live at freellmapi.co →

  • $19/year or $49 once, lifetime. Stripe checkout; cancel anytime, self-serve.
  • One fla_ key covers every router you run: desktop, homelab, Raspberry Pi.
  • Activate in the dashboard under Premium; cancel or manage billing self-serve at freellmapi.co/manage.
  • The router itself stays MIT-licensed and fully free, forever. Premium is only the live feed, and it's what funds the daily model testing and catalog maintenance that keeps the catalog working.

The catalog server never sees your prompts, completions, or provider keys — the router stays fully self-hosted either way.

Using the API

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

resp = client.chat.completions.create(
    model="auto",  # let the router pick; or "auto:fast", "auto:smart", a profile, or a model id
    messages=[{"role": "user", "content": "Summarise the fall of Rome in one sentence."}],
)
print(resp.choices[0].message.content)
print("Routed via:", resp.headers.get("x-routed-via"))

Streaming, the auto:* routing strategies, tool calling, vision input, Gemini Google Search grounding, embeddings, and the Anthropic Messages surface — with curl and Python examples for each — are all in docs/api.md. Every response carries an X-Routed-Via: <platform>/<model> header so you can see which provider actually served it.

Screenshots

Models

Pick a routing strategy and watch the monthly token budget fill across the whole provider fleet. Every model shows live reliability, speed, and intelligence scores — the order below is how requests route right now.

Models page

Keys

Manage provider credentials and grab the unified API key your apps connect with. Each key shows a status dot and when it was last health-checked.

Keys page

Playground

Send a chat completion through the router and see which provider served it, with the model ID and latency printed right on the message. Attach files by button, drag-and-drop, or paste: images (PNG/JPEG/WebP/GIF) are downscaled in the browser and sent as image content parts to a vision-capable model, and text files (TXT/MD/CSV/JSON/LOG) are inlined into the prompt as fenced blocks.

Playground page

Analytics

Request volume, success rate, tokens in and out, average latency, and per-provider breakdowns over 24h / 7d / 30d / 90d windows.

Analytics page

How it works

One request in, the best free model out — the fallback chain with live scores, cooldowns, and quota tracking

One request in, the best free model out: the router picks the highest-priority model with a healthy key that's under all its rate limits, decrypts the key in memory, and calls the provider — on a 429/5xx it cools that key down and retries the next model in your chain. The component walkthrough, routing internals, and operational details live in docs/architecture.md.

Limitations

Stacking free tiers has real trade-offs: no frontier models, variable latency, no SLA — and the effective intelligence of the endpoint dips late in the day as top models hit their daily caps, then resets at UTC midnight. Read the honest list in docs/architecture.md#limitations before building anything real on this.

Contributing

Contributors very welcome! See CONTRIBUTING.md for the dev loop, PR expectations, and the policy on AI/LLM-assisted contributions (short version: welcome, same quality bar as any other PR). Good first PRs:

  • Add a provider — copy server/src/providers/openai-compat.ts as a template, wire it into server/src/providers/index.ts, seed its models in server/src/db/index.ts, add a test in server/src/__tests__/providers/.
  • Add an endpoint — moderations and other OpenAI-compatible surfaces. The provider base class can grow new methods; adapters declare which they support.
  • Improve the router — cost-aware routing (cheapest-healthy-fastest tradeoffs), better latency-weighted priority, regional pinning.
  • Dashboard polish — charts on the Analytics page, key rotation UX, batch import of keys from .env.
  • Docs — more examples, client library snippets for Go/Rust/etc., a deployment recipe for Docker or Fly.

npm install && npm run dev gets you the server on :3001 and the dashboard on :5173, both with HMR. For a repeatable setup, use ./scripts/dev-bootstrap.sh on Bash or .\scripts\dev-bootstrap.ps1 on PowerShell; each preserves an existing .env. PRs should include a test, keep the existing suite green (npm test), and match the .editorconfig / tsconfig defaults already in the repo. Database migration workflow and the full contributor loop are in CONTRIBUTING.md.

Contributors

@moaaz12-web @lukasulc @VinhPhamAI @deadc @zhangyu1324 @chongjiazhen @vjsai @long2ice @sadesguy @hodlmybeer69-bit @phoenixikkifullstack @jtbrennan-git @praveenkumarpranjal @nordbyte @mybropro @danscMax @jhash @JammyJames1234 @coffcoe @Sumit4codes @meliani @thedavidweng @bharvey42

Recent activity

commits and pull requests

Recent open issues

view all

Discussions

all 7

Releases and announcements

12 total
  1. v0.6.8v0.6.8Aug 5, 2026219 downloads

    ## The desktop app comes to Linux This release ships native Linux desktop builds for the first time — an AppImage, a `.deb`, and a `.tar.xz`, all x64, attached below alongside the Windows and macOS installers (#739, thanks @UrbsKali). The AppImage carries `latest-linux.yml` update metadata like the other platforms. ## Routing - **Explore unmeasured models** (#731). A new model in your chain used to starve: with no reliability samples it never scored high enough to be picked, so it never earned any samples either. An opt-in toggle — tucked at the bottom of the Routing strategy card, off by default — gives models with fewer than 5 recent samples a 10% chance to be tried first so they build a track record. Manual mode is unaffected. - **Groundwork for community reliability priors** (#744). The scoring engine can now seed a brand-new model's reliability from community-sourced counts instead of a coin flip. It is entirely inert for now — nothing uploads or downloads anything, priors are capped so your own traffic always outweighs them, and the whole path sits behind a default-off setting until the data-sharing design lands. ## Models - Pick a model's capability tier — Frontier / La

  2. v0.6.7v0.6.7Aug 3, 2026614 downloads

    ## The dashboard loads again over plain HTTP If you reach FreeLLMAPI at `http://<ip>:<port>` — a Docker install on your LAN, a Raspberry Pi, a home server — v0.6.6 served you a blank page. This release fixes it, and it is the reason to upgrade (#682, #687, #734, reported by @EntropyEngineer). An origin without TLS is not a "secure context", and v0.6.6's CSP hardening did not account for that: - `upgrade-insecure-requests` was rewriting `/assets/*` to `https://` on an origin with no TLS, so every script and stylesheet failed with `ERR_SSL_PROTOCOL_ERROR`. - The inline theme bootstrap — the script that stops dark-mode users seeing a white flash — was blocked outright. It is now allowed by hash, and a test recomputes that hash from the real `index.html` so it cannot silently drift. - `Cross-Origin-Opener-Policy` and `Origin-Agent-Cluster` were being sent to origins that discard them and log an error instead. They now go out only over TLS or loopback, and come back automatically behind an HTTPS reverse proxy that forwards `X-Forwarded-Proto`. - Copy buttons had no working clipboard: `navigator.clipboard` does not exist on an insecure origin, so they either threw or silently did noth

  3. v0.6.6v0.6.6Jul 30, 20261.1K downloads

    ## Docker `:latest` now follows releases If you run FreeLLMAPI in Docker, this is the one to read. `:latest` was being retagged on every push to `main`, so pulling it gave you unreleased code instead of the newest release. It now follows release tags only (#679, reported in discussion #533). From this release onward, `:latest` and `:v0.6.6` are the same image: ```bash docker pull ghcr.io/tashfeenahmed/freellmapi:latest ``` Main builds keep their own `main` and `sha-<sha>` tags, so if you were deliberately tracking the development stream, use those. ## Admin surface hardening Thanks to @s-uryansh for this work (#498): - Stricter CSP via Helmet. - Per-IP rate limiting on `/api`, tunable with `ADMIN_RATE_LIMIT_RPM` (`0` disables). Key export gets its own tighter bucket, since it is the one admin endpoint that verifies a password. - Exporting your API keys now asks for your dashboard password. It is asked as a second step, after you choose what to export, rather than up front. - `X-Forwarded-For` is ignored unless `trust proxy` is enabled. - Production 5xx responses are sanitised, with full detail still going to the server log. ## Translations - zh-CN quality pass (#669). - zh

  4. v0.6.5v0.6.5Jul 29, 2026371 downloads

    The Agents page now identifies each supported coding agent by its own logo rather than a drawn approximation. ## Improvements - **Every supported coding agent is shown with its official brand mark, in that brand's own colours.** Where a project publishes separate artwork for light and dark backgrounds, the matching variant is used on each theme, so the marks stay legible in both. Crush, which publishes no vector mark, keeps a brand-tinted lettermark No configuration changes and no database migrations; upgrading is a straight image or app replacement. Full diff: https://github.com/tashfeenahmed/freellmapi/compare/v0.6.4...v0.6.5 > **macOS note:** the DMG is not notarized — on first launch use right-click → Open (or `xattr -d com.apple.quarantine /Applications/FreeLLMAPI.app`).

  5. v0.6.4v0.6.4Jul 29, 202695 downloads

    Custom relay endpoints that serve the same model id are now tracked independently. Previously every custom relay stored its models under a single identity, so registering a model id that another endpoint already served rebound the existing row — one enable flag, one set of scores, one cooldown shared between unrelated endpoints. ## Improvements - **Each custom endpoint now keeps its own model row** for a given model id, with independent enable state, reliability and speed scores, cooldowns and retirement. A relay that starts failing no longer demotes, cools down or disables the same model id on a different relay (#651, #619) - **Requesting a model by its plain id now reaches every endpoint that serves it**, so failover across relays works by default. When two endpoints serve the same id, the dashboard shows each endpoint's URL alongside the model and offers per-endpoint ids for pinning one specific endpoint - **A model requested by its exact id is always tried first.** Models matched only by a similar display name are used as fallback afterwards, never ahead of an exact match - **Setups with a single custom endpoint are unaffected** — no new fields to fill in, no naming changes,

Code frequency

additions and deletions
+79.2K-79.2KWeek of 2026-04-19: +20,854 linesWeek of 2026-04-19: -205 linesWeek of 2026-04-26: +1,136 linesWeek of 2026-04-26: -655 linesWeek of 2026-05-03: +44 linesWeek of 2026-05-03: -18 linesWeek of 2026-05-10: +73 linesWeek of 2026-05-10: -3 linesWeek of 2026-05-17: +1,778 linesWeek of 2026-05-17: -574 linesWeek of 2026-05-24: +1,399 linesWeek of 2026-05-24: -152 linesWeek of 2026-05-31: +22,065 linesWeek of 2026-05-31: -3,872 linesWeek of 2026-06-07: +12,060 linesWeek of 2026-06-07: -5,476 linesWeek of 2026-06-14: +5,975 linesWeek of 2026-06-14: -867 linesWeek of 2026-06-21: +9,525 linesWeek of 2026-06-21: -480 linesWeek of 2026-06-28: +3,995 linesWeek of 2026-06-28: -283 linesWeek of 2026-07-05: +16,629 linesWeek of 2026-07-05: -3,789 linesWeek of 2026-07-12: +1,075 linesWeek of 2026-07-12: -156 linesWeek of 2026-07-19: +3,873 linesWeek of 2026-07-19: -296 linesWeek of 2026-07-26: +79,247 linesWeek of 2026-07-26: -4,483 linesWeek of 2026-08-02: +18,370 linesWeek of 2026-08-02: -7,417 linesApr 19, 2026Aug 2, 2026
+198.1K lines added, -28.7K removed over the last year.

Commits per week

last 52 weeks
790Week of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 9 commitsWeek of 2026-04-26: 14 commitsWeek of 2026-05-03: 3 commitsWeek of 2026-05-10: 2 commitsWeek of 2026-05-17: 20 commitsWeek of 2026-05-24: 18 commitsWeek of 2026-05-31: 79 commitsWeek of 2026-06-07: 52 commitsWeek of 2026-06-14: 18 commitsWeek of 2026-06-21: 28 commitsWeek of 2026-06-28: 28 commitsWeek of 2026-07-05: 47 commitsWeek of 2026-07-12: 8 commitsWeek of 2026-07-19: 11 commitsWeek of 2026-07-26: 68 commitsWeek of 2026-08-02: 32 commitsAug 10, 2025Aug 2, 2026
437 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 8 commitsSun 1:00 — 9 commitsSun 2:00 — 1 commitsSun 3:00 — 1 commitsSun 4:00 — 2 commitsSun 5:00 — 0 commitsSun 6:00 — 1 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 8 commitsSun 10:00 — 1 commitsSun 11:00 — 3 commitsSun 12:00 — 6 commitsSun 13:00 — 7 commitsSun 14:00 — 7 commitsSun 15:00 — 4 commitsSun 16:00 — 4 commitsSun 17:00 — 7 commitsSun 18:00 — 4 commitsSun 19:00 — 2 commitsSun 20:00 — 1 commitsSun 21:00 — 2 commitsSun 22:00 — 1 commitsSun 23:00 — 4 commitsMon 0:00 — 2 commitsMon 1:00 — 4 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 2 commitsMon 5:00 — 1 commitsMon 6:00 — 0 commitsMon 7:00 — 2 commitsMon 8:00 — 2 commitsMon 9:00 — 0 commitsMon 10:00 — 1 commitsMon 11:00 — 3 commitsMon 12:00 — 4 commitsMon 13:00 — 3 commitsMon 14:00 — 3 commitsMon 15:00 — 2 commitsMon 16:00 — 3 commitsMon 17:00 — 2 commitsMon 18:00 — 0 commitsMon 19:00 — 3 commitsMon 20:00 — 0 commitsMon 21:00 — 4 commitsMon 22:00 — 4 commitsMon 23:00 — 3 commitsTue 0:00 — 2 commitsTue 1:00 — 6 commitsTue 2:00 — 1 commitsTue 3:00 — 0 commitsTue 4:00 — 4 commitsTue 5:00 — 0 commitsTue 6:00 — 1 commitsTue 7:00 — 0 commitsTue 8:00 — 0 commitsTue 9:00 — 0 commitsTue 10:00 — 0 commitsTue 11:00 — 0 commitsTue 12:00 — 3 commitsTue 13:00 — 8 commitsTue 14:00 — 8 commitsTue 15:00 — 3 commitsTue 16:00 — 9 commitsTue 17:00 — 0 commitsTue 18:00 — 1 commitsTue 19:00 — 2 commitsTue 20:00 — 5 commitsTue 21:00 — 3 commitsTue 22:00 — 4 commitsTue 23:00 — 7 commitsWed 0:00 — 2 commitsWed 1:00 — 5 commitsWed 2:00 — 0 commitsWed 3:00 — 4 commitsWed 4:00 — 1 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 5 commitsWed 9:00 — 3 commitsWed 10:00 — 3 commitsWed 11:00 — 2 commitsWed 12:00 — 2 commitsWed 13:00 — 2 commitsWed 14:00 — 7 commitsWed 15:00 — 2 commitsWed 16:00 — 3 commitsWed 17:00 — 0 commitsWed 18:00 — 1 commitsWed 19:00 — 0 commitsWed 20:00 — 5 commitsWed 21:00 — 6 commitsWed 22:00 — 0 commitsWed 23:00 — 2 commitsThu 0:00 — 5 commitsThu 1:00 — 6 commitsThu 2:00 — 2 commitsThu 3:00 — 2 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 0 commitsThu 9:00 — 0 commitsThu 10:00 — 5 commitsThu 11:00 — 3 commitsThu 12:00 — 8 commitsThu 13:00 — 3 commitsThu 14:00 — 3 commitsThu 15:00 — 8 commitsThu 16:00 — 5 commitsThu 17:00 — 0 commitsThu 18:00 — 1 commitsThu 19:00 — 0 commitsThu 20:00 — 0 commitsThu 21:00 — 1 commitsThu 22:00 — 0 commitsThu 23:00 — 1 commitsFri 0:00 — 2 commitsFri 1:00 — 0 commitsFri 2:00 — 1 commitsFri 3:00 — 0 commitsFri 4:00 — 2 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 2 commitsFri 8:00 — 1 commitsFri 9:00 — 0 commitsFri 10:00 — 7 commitsFri 11:00 — 4 commitsFri 12:00 — 3 commitsFri 13:00 — 5 commitsFri 14:00 — 9 commitsFri 15:00 — 8 commitsFri 16:00 — 7 commitsFri 17:00 — 0 commitsFri 18:00 — 1 commitsFri 19:00 — 2 commitsFri 20:00 — 5 commitsFri 21:00 — 0 commitsFri 22:00 — 2 commitsFri 23:00 — 4 commitsSat 0:00 — 8 commitsSat 1:00 — 1 commitsSat 2:00 — 1 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 5 commitsSat 11:00 — 2 commitsSat 12:00 — 1 commitsSat 13:00 — 12 commitsSat 14:00 — 4 commitsSat 15:00 — 6 commitsSat 16:00 — 5 commitsSat 17:00 — 1 commitsSat 18:00 — 2 commitsSat 19:00 — 2 commitsSat 20:00 — 1 commitsSat 21:00 — 5 commitsSat 22:00 — 2 commitsSat 23:00 — 8 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits393 (78%)
Community commits109 (22%)

502 commits in total over the last year.

DateListRankStars gained
Jun 25, 2026daily#14+21
May 21, 2026daily#16+56
May 20, 2026daily#15+54
  • freeCodeCamp/freeCodeCamp

    freeCodeCamp.org's open-source codebase and curriculum. Learn math, programming, and computer science for free.

    453.6K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    385.5K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • anomalyco/opencode

    The open source coding agent.

    194.7K stars · TypeScript