Trending repositories: local-llm

16 tracked repositories tagged with local-llm, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

16 of 16 repositories

  • Alishahryar1/free-claude-code

    Use Claude Code, Codex, VSCode, Pi, and OpenCode (and 6 other harnesses) for free (1.3B+ free tokens) from your terminal, app, IDE, or phone, and now from the browser with native browser sessions (multi-harness + multi-model) like OpenClaw (voice supported + ToS friendly)

    AI summary: A local API gateway that intercepts and routes requests from AI coding agents to alternative local or cloud LLMs.

    56,450developer-toolsPythonOther
  • HKUDS/nanobot

    Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps

    AI summary: An ultra-lightweight, self-hosted Python framework for running persistent personal AI agents with native MCP support.

    48,762ai-mlPythonMIT
  • JustVugg/colibri

    Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

    AI summary: A pure C inference engine for running massive MoE models locally on consumer hardware using memory tiering.

    39,323ai-mlCApache-2.0
  • agentscope-ai/QwenPaw

    Your Personal AI Assistant; easy to install, deploy on your own machine or on the cloud; supports multiple chat apps with easily extensible capabilities.

    AI summary: A highly extensible personal AI assistant featuring three-layer memory and multi-channel integration.

    35,440ai-mlTypeScriptApache-2.0
  • garrytan/gbrain

    Garry's Opinionated OpenClaw/Hermes Agent Brain

    AI summary: A highly customized local knowledge brain daemon for executing the OpenClaw and Hermes AI agents.

    30,492productivityTypeScriptMIT
  • tonhowtf/omniget

    Udemy & Hotmart course downloader, YouTube downloader (yt-dlp GUI, 1,800+ sites) + desktop app for AI agents: Claude Code, Codex, Gemini CLI, Ollama. Permissions, undo, jobs, loops until tests pass, MCP server, 156 tools, course player. Free and open source for Windows, macOS and Linux. No terminal. Your files stay on your computer.

    AI summary: A universal, cross-platform app for downloading media from over 1,800 sites without using the terminal.

    14,302productivityRustGPL-3.0
  • LearningCircuit/local-deep-research

    ~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.

    AI summary: Fully local, agentic AI assistant for deep web research with accurate citations on consumer hardware.

    9,149ai-mlPythonMIT
  • MakazhanAlpamys/Soup

    Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

    AI summary: A lightning-fast, C++ based LLM fine-tuning CLI that utilizes layer streaming.

    8,070ai-mlPythonApache-2.0
  • OpenCoworkAI/open-codesign

    Open-source Claude Design alternative. One-click import your Claude Code / Codex API key. Prompt → prototype / slides / PDF. Multi-model (Claude, GPT, Gemini, Kimi, GLM, Ollama). BYOK, local-first, MIT.

    AI summary: A desktop application that utilizes local AI to rapidly generate, edit, and export UI components and landing pages.

    7,997developer-toolsTypeScriptMIT
  • open-multi-agent/open-multi-agent

    Self-hosted TypeScript agent runtime with durable approvals and verifiable run records. Own it, approve it, audit it.

    AI summary: TypeScript orchestration framework that dynamically plans and executes tasks using multiple AI agents locally.

    6,972ai-mlTypeScriptMIT
  • Osmantic/ODS

    ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

    AI summary: An automated deployment system that wires together Ollama, Open WebUI, and workflow tools to create a private AI server.

    6,969ai-mlPythonApache-2.0
  • magnitudedev/magnitude

    Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or just a CPU.

    AI summary: An open-source inference server that seamlessly connects AI coding agents to locally hosted models

    6,274developer-toolsRustApache-2.0
  • PrismML-Eng/Bonsai-demo

    Bonsai Demo

    AI summary: A comprehensive demo repository for running the highly efficient 1-bit Bonsai and Ternary-Bonsai language models locally.

    3,266ai-mlShellApache-2.0
  • youssofal/MTPLX

    The fastest way to run Qwen 3.8 Flash Next, Qwen 3.8 27B and Ternary Bonsai 2 27B on a Mac: 125 tok/s in OpenCode on an M5 Max, and a 27B model on 16 GB Macs. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.

    AI summary: A native Mac app and CLI tool for running local language models significantly faster utilizing multi-token prediction.

    2,507ai-mlPythonApache-2.0
  • jamesob/local-llm

    Everything I know about running LLMs locally

    AI summary: An opinionated, human-written guide for building local AI hardware and running LLMs using Docker.

    1,856infrastructureShell
  • localgpt-app/localgpt

    Local AI assistant, dreaming explorable worlds.

    AI summary: A local-first AI assistant and world-building engine powered by Rust and Bevy.

    1,123ai-mlRustApache-2.0