Mesh-LLM/mesh-llmPublic

Distributed AI/LLM for the people. Share compute privately or publicly to power your agents and chat.

AI summary: Pools GPUs across multiple machines into a single OpenAI-compatible API for distributed LLM inference.

Stars
3.5K
+4 today
Forks
429
Watchers
17
Open issues
106
Open PRs
6
Contributors
~41
Commits
2.7K
Branches
511

RustApache-2.0Created Feb 11, 2026Last push todayLatest release v0.77.0+19 stars this week+121 this month

Quick answers

What is mesh-llm?
Pools GPUs across multiple machines into a single OpenAI-compatible API for distributed LLM inference.
What does mesh-llm do?
Mesh LLM is a distributed inference engine that aggregates GPU and memory resources across multiple physical machines to run large language models. It exposes this pooled hardware as a single, unified OpenAI-compatible API endpoint on localhost. The system handles intelligent routing, automatically deciding whether to run a model locally, forward the request to a peer node, or split the model across machines using 'Skippy stage splits'. This allows users to run massive models that exceed the VRAM capacity of any single machine in their network. It drastically simplifies the deployment of complex, distributed AI infrastructure.
Who is mesh-llm for?
AI developers and researchers who need to run large models but lack single machines with massive VRAM. It is also ideal for homelab enthusiasts and small startups looking to build cost-effective, distributed AI infrastructure.
How do I get started with mesh-llm?
curl -fsSL https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.sh | bash
How popular is mesh-llm on GitHub?
Mesh-LLM/mesh-llm has 3,476 stars and 429 forks on GitHub, and gained 19 stars in the last 7 days.
What license does mesh-llm use?
Mesh-LLM/mesh-llm is released under the Apache-2.0 license.

Star history

since Jul 28, 2026
01K2K3KJul 2026Aug 2026Sep 2026Oct 2026
3.5K stars as of Oct 4, 2026. Measured daily since Jul 28, 2026; GitHub no longer exposes earlier star timestamps.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugSepMonWedFri2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 3 commits2026-02-12: 26 commits2026-02-13: 11 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 13 commits2026-02-17: 12 commits2026-02-18: 12 commits2026-02-19: 19 commits2026-02-20: 20 commits2026-02-21: 3 commits2026-02-22: 0 commits2026-02-23: 24 commits2026-02-24: 23 commits2026-02-25: 21 commits2026-02-26: 19 commits2026-02-27: 0 commits2026-02-28: 10 commits2026-03-01: 10 commits2026-03-02: 2 commits2026-03-03: 3 commits2026-03-04: 3 commits2026-03-05: 19 commits2026-03-06: 6 commits2026-03-07: 3 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 9 commits2026-03-11: 14 commits2026-03-12: 15 commits2026-03-13: 7 commits2026-03-14: 10 commits2026-03-15: 7 commits2026-03-16: 6 commits2026-03-17: 7 commits2026-03-18: 6 commits2026-03-19: 19 commits2026-03-20: 15 commits2026-03-21: 4 commits2026-03-22: 4 commits2026-03-23: 34 commits2026-03-24: 25 commits2026-03-25: 12 commits2026-03-26: 6 commits2026-03-27: 35 commits2026-03-28: 29 commits2026-03-29: 26 commits2026-03-30: 44 commits2026-03-31: 12 commits2026-04-01: 28 commits2026-04-02: 49 commits2026-04-03: 63 commits2026-04-04: 60 commits2026-04-05: 10 commits2026-04-06: 28 commits2026-04-07: 18 commits2026-04-08: 26 commits2026-04-09: 46 commits2026-04-10: 35 commits2026-04-11: 25 commits2026-04-12: 15 commits2026-04-13: 32 commits2026-04-14: 23 commits2026-04-15: 14 commits2026-04-16: 30 commits2026-04-17: 15 commits2026-04-18: 35 commits2026-04-19: 12 commits2026-04-20: 8 commits2026-04-21: 9 commits2026-04-22: 9 commits2026-04-23: 4 commits2026-04-24: 5 commits2026-04-25: 7 commits2026-04-26: 10 commits2026-04-27: 2 commits2026-04-28: 3 commits2026-04-29: 3 commits2026-04-30: 2 commits2026-05-01: 5 commits2026-05-02: 5 commits2026-05-03: 5 commits2026-05-04: 2 commits2026-05-05: 7 commits2026-05-06: 13 commits2026-05-07: 10 commits2026-05-08: 12 commits2026-05-09: 12 commits2026-05-10: 13 commits2026-05-11: 9 commits2026-05-12: 9 commits2026-05-13: 10 commits2026-05-14: 5 commits2026-05-15: 1 commit2026-05-16: 4 commits2026-05-17: 3 commits2026-05-18: 5 commits2026-05-19: 5 commits2026-05-20: 11 commits2026-05-21: 11 commits2026-05-22: 8 commits2026-05-23: 6 commits2026-05-24: 11 commits2026-05-25: 12 commits2026-05-26: 12 commits2026-05-27: 17 commits2026-05-28: 10 commits2026-05-29: 21 commits2026-05-30: 12 commits2026-05-31: 7 commits2026-06-01: 3 commits2026-06-02: 7 commits2026-06-03: 6 commits2026-06-04: 7 commits2026-06-05: 10 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 2 commits2026-06-11: 4 commits2026-06-12: 15 commits2026-06-13: 11 commits2026-06-14: 7 commits2026-06-15: 1 commit2026-06-16: 2 commits2026-06-17: 4 commits2026-06-18: 6 commits2026-06-19: 5 commits2026-06-20: 2 commits2026-06-21: 2 commits2026-06-22: 4 commits2026-06-23: 3 commits2026-06-24: 3 commits2026-06-25: 2 commits2026-06-26: 4 commits2026-06-27: 1 commit2026-06-28: 5 commits2026-06-29: 13 commits2026-06-30: 7 commits2026-07-01: 2 commits2026-07-02: 0 commits2026-07-03: 2 commits2026-07-04: 1 commit2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 1 commit2026-07-08: 0 commits2026-07-09: 0 commits2026-07-10: 2 commits2026-07-11: 4 commits2026-07-12: 3 commits2026-07-13: 2 commits2026-07-14: 11 commits2026-07-15: 10 commits2026-07-16: 2 commits2026-07-17: 6 commits2026-07-18: 5 commits2026-07-19: 3 commits2026-07-20: 2 commits2026-07-21: 8 commits2026-07-22: 7 commits2026-07-23: 2 commits2026-07-24: 7 commits2026-07-25: 3 commits2026-07-26: 5 commits2026-07-27: 8 commits2026-07-28: 8 commits2026-07-29: 8 commits2026-07-30: 10 commits2026-07-31: 0 commits2026-08-01: 2 commits2026-08-02: 8 commits2026-08-03: 2 commits2026-08-04: 6 commits2026-08-05: 8 commits2026-08-06: 2 commits2026-08-07: 3 commits2026-08-08: 6 commits2026-08-09: 8 commits2026-08-10: 8 commits2026-08-11: 8 commits2026-08-12: 13 commits2026-08-13: 12 commits2026-08-14: 24 commits2026-08-15: 6 commits2026-08-16: 0 commits2026-08-17: 7 commits2026-08-18: 12 commits2026-08-19: 7 commits2026-08-20: 9 commits2026-08-21: 7 commits2026-08-22: 2 commits2026-08-23: 1 commit2026-08-24: 15 commits2026-08-25: 3 commits2026-08-26: 8 commits2026-08-27: 6 commits2026-08-28: 14 commits2026-08-29: 8 commits2026-08-30: 25 commits2026-08-31: 7 commits2026-09-01: 19 commits2026-09-02: 11 commits2026-09-03: 10 commits2026-09-04: 16 commits2026-09-05: 7 commits2026-09-06: 1 commit2026-09-07: 2 commits2026-09-08: 14 commits2026-09-09: 12 commits2026-09-10: 11 commits2026-09-11: 22 commits2026-09-12: 20 commits2026-09-13: 8 commits2026-09-14: 10 commits2026-09-15: 16 commits2026-09-16: 17 commits2026-09-17: 3 commits2026-09-18: 6 commits2026-09-19: 9 commits2026-09-20: 17 commits2026-09-21: 9 commits2026-09-22: 21 commits2026-09-23: 0 commits2026-09-24: 0 commits2026-09-25: 0 commits2026-09-26: 0 commits
2,273 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Very active

    2,273 commits in 52 weeks

  • Permissive license

    Apache-2.0

What mesh-llm does

Mesh LLM is a distributed inference engine that aggregates GPU and memory resources across multiple physical machines to run large language models. It exposes this pooled hardware as a single, unified OpenAI-compatible API endpoint on localhost. The system handles intelligent routing, automatically deciding whether to run a model locally, forward the request to a peer node, or split the model across machines using 'Skippy stage splits'. This allows users to run massive models that exceed the VRAM capacity of any single machine in their network. It drastically simplifies the deployment of complex, distributed AI infrastructure.

AI developers and researchers who need to run large models but lack single machines with massive VRAM. It is also ideal for homelab enthusiasts and small startups looking to build cost-effective, distributed AI infrastructure.

  • Resource pooling: Aggregates VRAM and compute from multiple disparate nodes into a unified mesh.
  • OpenAI API compatibility: Exposes a standard API (`http://localhost:9337/v1`) for seamless integration with existing tools.
  • Intelligent routing: Automatically decides the most efficient node or nodes to handle specific inference requests.
  • Distributed inference (Skippy splits): Capable of splitting large models across multiple machines when they don't fit on one.
  • Dynamic scaling: Allows nodes to be added to the mesh dynamically to increase total capacity over time.

Where teams use it

Running massive LLMs locally

Deploying models like Llama-3-70B by pooling the VRAM of several smaller consumer GPUs across a local network.

Cost-effective AI infrastructure

Utilizing existing, fragmented hardware resources instead of renting expensive, high-end cloud GPUs.

High-availability inference

Creating a robust local API endpoint that can route around node failures or heavy load.

Collaborative compute clusters

Allowing small teams to pool their individual workstations into a shared AI inference server.

Getting started: curl -fsSL https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.sh | bash

README

main branch

Mesh LLM

Mesh LLM web console

Mesh LLM pools GPUs and memory across machines and exposes the result as one OpenAI-compatible API at http://localhost:9337/v1. Start one node, add more nodes later, and let the mesh decide whether a model runs locally, routes to a peer, or uses Skippy stage splits for models that are too large for one box.

Quick start

Install the latest release executable:

curl -fsSL https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.sh | bash

On Windows, use PowerShell:

irm https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.ps1 | iex

Install the Apple Silicon Homebrew formula with brew install Mesh-LLM/tap/mesh-llm. Versioned formulas, Ubuntu and Arch packages, checksums, SBOMs, and OCI images are produced by the public Mesh-LLM/mesh-packaging repository. See the platform install guides for the supported package matrix and install commands.

Finish setup:

mesh-llm setup

On Windows PowerShell, use mesh-llm.exe setup. (For native Windows notes, and for the optional WSL2 setup and multi-node LAN clustering, see the Windows & WSL2 Troubleshooting Guide.)

To remove an executable install later, preview the cleanup first:

mesh-llm uninstall --dry-run
mesh-llm uninstall --yes

Uninstall preserves ~/.mesh-llm configuration and identity data unless you explicitly pass --purge-config.

Join the public mesh and start serving:

mesh-llm serve --auto

That command chooses a backend flavor, downloads a suitable model if needed, joins the best discovered public mesh, starts the local API on port 9337, and starts the web console on port 3131.

Check available models:

curl -s http://localhost:9337/v1/models | jq '.data[].id'

Send an OpenAI-compatible request:

curl http://localhost:9337/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"GLM-4.7-Flash-Q4_K_M","messages":[{"role":"user","content":"hello"}]}'

For server deployments, add --headless to hide the web UI while keeping the management API on the --console port:

mesh-llm serve --auto --headless

Pick the workflow you need

Goal Command Full guide
Try the public mesh mesh-llm serve --auto docs/MESHES.md
Start a private mesh mesh-llm serve --model Qwen3-8B-Q4_K_M docs/MESHES.md
Serve one model without mesh networking (debugging/dev) mesh-llm serve --local-model-only --gguf /models/model.gguf OpenAI API defaults to 127.0.0.1:9337 (--port and --listen-all change it)
Publish your own mesh mesh-llm serve --model Qwen3-8B-Q4_K_M --publish docs/MESHES.md
Join by invite token mesh-llm serve --join <token> docs/MESHES.md
Run an API-only client mesh-llm client --auto docs/MESHES.md
Run a big model with splits mesh-llm serve --model hf://meshllm/<repo>@<rev> --split docs/SKIPPY_SPLITS.md
Run DeepSeek V4 and other ds4 models with DwarfStar mesh-llm plugins install Mesh-LLM/ds4-plugin Setup and configuration
Attach a Flash-MoE SSD backend mesh-llm serve with [[plugin]] name = "flash-moe" docs/plugins/flash-moe.md
Fan out one prompt to every model in the mesh curl ... -d '{"model":"mesh", ...}' docs/design/MOA_GATEWAY.md
Use Goose, OpenCode, Claude Code, or Pi mesh-llm goose, mesh-llm opencode, mesh-llm claude, mesh-llm pi docs/AGENTS.md
Build or contribute just build CONTRIBUTING.md

How the mesh works

  • Single-machine fit first. If one node can host the full model, it serves the model locally without stage traffic.
  • Mesh routing. Every node exposes the same /v1 API. Requests are routed by the model field to the peer that can serve that model.
  • Encrypted peer transport. QUIC end-to-end encrypts traffic between Mesh nodes, including inference requests, responses, and split-model activations. Iroh relays forward encrypted packets without reading their payload.
  • Owner-control plane. Operator config and inventory actions use an additive mesh-llm-control/1 lane with explicit endpoint bootstrap, while public mesh join, gossip, routing, and inference stay on the public mesh plane for mixed-version compatibility.
  • Skippy stage splits. Large dense models can load as package-backed layer stages. The coordinator plans contiguous layer ranges, starts downstream stages first, waits for readiness, then publishes the stage-0 route.
  • Layer packages. Package repositories contain model-package.json plus GGUF fragments so peers fetch only the pieces needed for their assigned stage.
  • Public discovery. Published meshes advertise through Nostr discovery; private meshes stay invite-token based.

For a deeper operator guide, see docs/USAGE.md. For every CLI command and switch, see docs/CLI.md.

Local model-only serving

Debugging and development endpoint. --local-model-only wires the OpenAI frontend directly to one local Skippy instance and bypasses mesh networking entirely. The supported path for OpenAI clients in production — including tool calls — is the regular mesh inference port, which includes the host-runtime normalizer. Use --local-model-only when you want a lightweight single-process setup without joining a mesh, not as a drop-in for the full serving path.

Use the direct topology when a process should expose one complete local model through the OpenAI API without becoming a mesh node:

mesh-llm serve \
  --local-model-only \
  --gguf /models/model.gguf \
  --port 9337

Use --gguf <absolute-path> to serve a specific GGUF file directly. --model <absolute-path> also works. Note that --model <catalog-name> does not work here — catalog and HF-ref resolution is a mesh function and is not available in the direct topology; the resulting error is not obvious. See the path requirements below.

This mode starts the OpenAI frontend and one local Skippy model runtime. It does not start QUIC, discovery, peer maintenance, split planning, plugins, release lookup, the web console, or the management API. Add --listen-all only when the OpenAI endpoint must bind beyond loopback. Startup fails if the complete model does not fit within detected local capacity (or --max-vram); it never falls back to distributed serving.

Values passed to --model, --gguf, and --mmproj must be absolute paths and must not be symlinks. If you downloaded a model through Hugging Face, resolve the cache symlink before passing the path: realpath <path-from-hf-cache>.

Mixture-of-Agents (model: "mesh") — experimental

⚠️ Experimental. The MoA gateway is new in this release. Behavior, routing heuristics, error shapes, and tuning knobs may change between versions while we tune it. Treat model: "mesh" as a preview feature rather than a stable production path; use a specific model id when you need stable semantics.

Send a request with "model": "mesh" and the proxy fans it out to every model available in the mesh in parallel, arbitrates their responses with deterministic logic, and returns one OpenAI-compatible reply. The arbiter runs in code (not as another model call) and only escalates to a reducer LLM on genuine conflict. Tool calls flow through the full pipeline.

curl http://localhost:9337/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"mesh","messages":[{"role":"user","content":"What is the capital of Japan?"}]}'

Requires at least two distinct models in the mesh. See docs/design/MOA_GATEWAY.md for the architecture, arbitration rules, and tuning knobs.

Supported model families

Mesh LLM's Skippy runtime tracks llama.cpp family parity with reviewed GGUF representatives. The current reviewed support set covers 72 P0/P1 family rows, with 89 certified rows in the full parity inventory, including Qwen, Llama, Gemma, Mistral, DeepSeek, GLM, MiniMax, Phi, Granite, Hunyuan, EXAONE, Cohere, Falcon, RWKV, and many others.

Split multimodal serving is certified for Qwen2-VL, Qwen3-VL, Qwen3-VL-MoE, HunyuanOCR/Hunyuan-VL, and DeepSeek-OCR using real GGUF plus projector fixtures. DeepSeek3 and EXAONE-MoE use package-backed stages because the full GGUFs are too large for the cheap local baseline.

See docs/skippy/FAMILY_STATUS.md for the full artifact, split, wire dtype, cache policy, and exception matrix. See docs/skippy/LLAMA_PARITY.md for the remaining llama.cpp parity queue.

Install and build notes

Tagged releases publish macOS bundles plus Linux CPU, Linux ARM64 CPU, Linux ARM64 CUDA, Linux CUDA, Linux CUDA Blackwell, Linux ROCm, Linux Vulkan, Windows CPU, Windows CUDA, Windows ROCm, and Windows Vulkan bundles. Metal is macOS-only. Every flavor is composed from the same backend-neutral host for its OS/architecture plus one versioned native runtime. The Linux ARM64 CPU artifact is mesh-llm-aarch64-unknown-linux-gnu.tar.gz; the Linux ARM64 CUDA artifact is mesh-llm-aarch64-unknown-linux-gnu-cuda.tar.gz. In install and release contexts, arm64 and aarch64 mean the same 64-bit ARM target. Portable archives work offline: the host discovers the adjacent native-runtimes/<runtime-id> tree before consulting the user cache.

Build from source with just:

git clone https://github.com/Mesh-LLM/mesh-llm
cd mesh-llm
just build

Source builds require just, cmake, Rust, sccache, the platform fast linker (mold on Linux and lld on macOS/Windows), and Node.js 24 + npm. Cargo uses these defaults from the checked-in configuration, including plain cargo build. See CONTRIBUTING.md for installation commands and linker fallback behavior. To exercise the release boundary locally, build the neutral host and one runtime:

just release-host-build
just release-runtime-build metal # or cpu, cuda, rocm, vulkan
MESH_LLM_NATIVE_RUNTIME_BUNDLE_DIR="$PWD/dist/native-runtimes" \
MESH_LLM_NATIVE_RUNTIME_CACHE_DIR="$(mktemp -d)" \
  ./target/release/mesh-llm runtime list

CUDA runtimes need nvcc, ROCm runtimes need ROCm/HIP, and Vulkan runtimes need Vulkan development files plus glslc. See docs/design/NATIVE_RUNTIMES.md for the manifest, discovery, and compatibility contract.

The shipped mesh-llm executable uses embedded release attestation for provenance and admission hardening only. It does not apply to SDK, XCFramework, or other native artifacts, and it is not a runtime integrity proof. Verify a stamped packaged executable with cargo run -p xtask -- release-attestation inspect --binary <path-to-packaged-mesh-llm> --public-key-file <release-signing-public-key.json>. A packaged release binary reports valid, an unstamped local or dev build reports missing, and a binary that changed after packaging reports invalid. Bare inspect --binary ... is only enough to classify an unstamped binary as missing; stamped binaries require --public-key-file and otherwise report invalid with an explicit error. Post-download mutation can flip a stamped binary to invalid, but default startup still allows it.

🪟 Windows & WSL2 Troubleshooting

Running natively on Windows (NVIDIA)

Native Windows CUDA works, including on CUDA 13.x drivers: GPU detection (mesh-llm gpus) and full-speed CUDA inference have been verified on driver 610.74 (CUDA UMD 13.3) with an RTX 4070 Ti. The earlier advice to switch to WSL2 when mesh-llm reported 0 GPUs on CUDA 13 drivers predates the cuda12-runtime compatibility fix (#1127) and no longer applies to current releases.

As of v0.76.0-rc8, three distribution/loading bugs still block the out-of-the-box native path. Until the fixes ship, this sequence works end to end:

  1. Install the prerelease — the stable v0.75.1 Windows bundles fail install.ps1 verification (#1510):

    irm https://raw.githubusercontent.com/Mesh-LLM/mesh-llm/main/install.ps1 -OutFile install.ps1
    .\install.ps1 -PreRelease
  2. Install the CUDA runtime from the product bundle — mesh-llm runtime install cuda finds no windows/x86_64 runtimes in the release manifest (#1511). Download mesh-llm-x86_64-pc-windows-msvc-cuda.zip for your installed version from the releases page, extract it, then:

    mesh-llm runtime install --bundle-dir "<extracted>\mesh-bundle" cuda
  3. Put the runtime's lib directory on PATH before serving — runtime DLLs currently fail to load with LoadLibraryExW error 126 (#1512):

    $env:PATH = "$env:LOCALAPPDATA\mesh-llm\native-runtimes\<version>\meshllm-native-runtime-windows-x86_64-cuda12\lib;" + $env:PATH
    mesh-llm serve --local-model-only --model "C:\path\to\model.gguf"

Good to know on native Windows:

  • --local-model-only requires an absolute path to a local .gguf file; catalog and Hugging Face refs are rejected in this mode.
  • Models download to %LOCALAPPDATA%\huggingface\hub (the Rust hf-hub convention) — a separate cache from the Python tools' ~\.cache\huggingface.
  • mesh-llm.exe is not code-signed yet, so SmartScreen may prompt when launching it manually.

Running under WSL2 (alternative)

WSL2 remains a solid alternative if you prefer the Linux CUDA runtimes or want the multi-node LAN setups described below.

1. Install CUDA 13.0 Toolkit inside WSL2

Inside your Ubuntu WSL2 terminal, install cuda-toolkit-13-0 to supply libcudart.so.13:

wget https://developer.download.nvidia.com/compute/cuda/repos/wsl-ubuntu/x86_64/cuda-wsl-ubuntu.pin
sudo mv cuda-wsl-ubuntu.pin /etc/apt/preferences.d/cuda-repository-pin-600
wget https://developer.download.nvidia.com/compute/cuda/13.0.0/local_installers/cuda-repo-wsl-ubuntu-13-0-local_13.0.0-1_amd64.deb
sudo dpkg -i cuda-repo-wsl-ubuntu-13-0-local_13.0.0-1_amd64.deb
sudo cp /var/cuda-repo-wsl-ubuntu-13-0-local/cuda-*-keyring.gpg /usr/share/keyrings/
sudo apt-get update && sudo apt-get -y install cuda-toolkit-13-0

echo 'export LD_LIBRARY_PATH=/usr/local/cuda-13.0/lib64:$LD_LIBRARY_PATH' >> ~/.bashrc
source ~/.bashrc
2. Enable Hyper-V & Windows Firewall for WSL2 Mirrored Mode

If you use WSL2 networkingMode=mirrored in %UserProfile%\.wslconfig, Windows 11 manages a separate Hyper-V VM Firewall that defaults to Block for inbound network traffic when third-party security software (e.g. Norton, McAfee) is present.

2.1 Configure Mirrored Networking (.wslconfig) Part 1

Create or edit C:\Users\<username>\.wslconfig on the Windows host:

notepad $env:USERPROFILE\.wslconfig
2.2 Configure Mirrored Networking (.wslconfig) Part 2
[wsl2]
networkingMode=mirrored
autoProxy=true
2.3 Restart WSL in Powershell

Restart WSL in PowerShell:

wsl --shutdown
2.4 Enable Windows Firewall Needs

To allow incoming LAN connections to the Web UI (3131) and P2P QUIC mesh transport (9337), run PowerShell as Administrator on the Windows host:

# 1. Allow Inbound Traffic through the WSL Hyper-V VM Firewall Container
$vmCreatorId = '{40E0AC32-46A5-438A-A0B2-2B479E8F2E90}'
Set-NetFirewallHyperVVMSetting -Name $vmCreatorId -DefaultInboundAction Allow

# 2. Allow MeshLLM Ports in Windows Defender Firewall
New-NetFirewallRule -DisplayName "MeshLLM TCP In" -Direction Inbound -Action Allow -Protocol TCP -LocalPort 3131,9337 -Profile Any
New-NetFirewallRule -DisplayName "MeshLLM UDP In" -Direction Inbound -Action Allow -Protocol UDP -LocalPort 9337,5353 -Profile Any
3. Match Model Paths for Direct LAN Reading

To ensure worker nodes load GGUF model shards directly off local NVMe/SSD storage without streaming tens of gigabytes over the network, ensure the --gguf file path string is identical across all nodes (or use symlinks/bind mounts):

# Example: Mount or symlink model path on worker nodes
sudo mkdir -p /mnt/d/models/
sudo mount --bind /path/to/local/fast/nvme/ /mnt/d/models/

# Launch Master and Worker with matching path strings
mesh-llm --llama-flavor cuda serve \
  --console 3131 \
  --gguf "/mnt/d/models/model.gguf" \
  --mesh-name "MainMesh" \
  --listen-all \
  --auto

Documentation hub

Doc Use it for
docs/MESHES.md Private meshes, public discovery, publishing, invite tokens, API-only clients
docs/SKIPPY_SPLITS.md Running big models with package-backed Skippy stage splits
docs/LAYER_PACKAGE_REPOS.md Contributing and publishing layer package repositories
docs/AGENTS.md Goose, Claude Code, OpenCode, Pi, curl, and blackboard
docs/EXO_COMPARISON.md Balanced comparison with Exo
docs/CLI.md Command reference and JSON automation
docs/USAGE.md Longer operational usage guide, runtime control, owner-control operator flows
docs/design/TESTING.md Testing playbook, mixed-version QA, remote deploy checks
docs/plugins/dwarfstar.md DwarfStar (ds4) alternative engine: setup on Apple Silicon
docs/plugins/flash-moe.md Optional Flash-MoE SSD expert streaming backend setup
docs/skippy/FAMILY_STATUS.md Certified Skippy model-family status
docs/specs/layer-package-repos.md Manifest and artifact format spec
docs/specs/mesh-setup-installer.md Installer/bootstrap and setup command behavior spec

CI infrastructure

Depot

Trusted main Linux jobs may use Depot's managed GitHub Actions runners through the checked-in runner policy. Pull requests remain on GitHub-hosted runners until the cache isolation and protected-workflow gates in ci/DEPOT_MIGRATION.md are proven. Hardware-qualified GPU tests remain on dedicated runners.

Community

Mesh LLM is experimental distributed-systems software. When you report bugs, include the command you ran, platform/backend flavor, /api/status output if available, and whether the node was private, published, or joined with --auto.

Support

Spiral

The Mesh LLM project thanks spiral.xyz for their support.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

143 total
  1. v0.77.0v0.77.0Sep 28, 20264.5K downloads

    ## [0.77.0] - 2026-09-28 ### Added #### skippy * Expand and streamline llama family certification by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1865 * Recover llama family certification through Inkling by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1880 * Add run-ahead verify windows by @danielwinterw in https://github.com/Mesh-LLM/mesh-llm/pull/1409 * Cut over to graph-derived stage filtering by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1927 * Bound exact-state recorder admission with two-class credits by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1936 * Add draft fallback for N-gram misses by @danielwinterw in https://github.com/Mesh-LLM/mesh-llm/pull/1410 * Serve non-chat GGUF workloads via OpenAI API by @IvGolovach in https://github.com/Mesh-LLM/mesh-llm/pull/1833 * Surface graph reuse counters to the host by @danielwinterw in https://github.com/Mesh-LLM/mesh-llm/pull/2016 * Replay extracted stage programs by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/2025 * Wire next-generation KV cache and publisher defaults by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1838 * Serve Laya decision models on /systemone by @danielwinterw in http

  2. v0.76.2v0.76.2Sep 14, 202625.8K downloads

    **Full Changelog**: https://github.com/Mesh-LLM/mesh-llm/compare/v0.76.1...v0.76.2

  3. v0.76.1v0.76.1Sep 12, 202625.4K downloads

    ## [0.76.1] - 2026-09-12 ### Added * Itemize the advertised capacity behind vram_bytes by @Virgile-pct in https://github.com/Mesh-LLM/mesh-llm/pull/1673 * Gpu_name_source beside gpu_name — how the value was captured, display unchanged by @StevenMih in https://github.com/Mesh-LLM/mesh-llm/pull/1679 * Diagnose direct-connect advertisement and peer paths by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1803 * Surface the advertised memory breakdown by @Virgile-pct in https://github.com/Mesh-LLM/mesh-llm/pull/1754 ### Fixed * Roll up the 0.76.1 bugfix branch into main by @ndizazzo in https://github.com/Mesh-LLM/mesh-llm/pull/1789 * Reject malformed tool definitions by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1761 * Link Apple OpenMP runtime by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1766 * Advertise externally observed public address by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1774 * Reject self-identity joins and add node key override by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1775 * Keep direct GGUF split loads local by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1776 * Report client mode from --client --auto console state by

  4. v0.76.0v0.76.0Sep 10, 20264.8K downloads

    ## [0.76.0] - 2026-09-10 ### Added #### Serving and inference * Durable KV prefix cache: agent prefixes survive eviction, restart, and cold nodes by @michaelneale in https://github.com/Mesh-LLM/mesh-llm/pull/1228 * skippy-quantize: compose-mtp — splice an MTP draft into a sharded target GGUF by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1439 * Load SafeTensors checkpoints directly by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1619 * Complete graph-derived v2 model splitting by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1672 * Iteration-level scheduler for concurrent staged serving (#1416) by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1420 * Route repeated prompts to workers with verified radix cache reuse by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1449 * Reuse warm prefixes across agent sessions by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1601 * Stream tool calls as they are generated instead of after the turn ends by @michaelneale in https://github.com/Mesh-LLM/mesh-llm/pull/1371 * Pass reasoning effort through Skippy templates by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1211 * Expose model-bound tokenizer capabilit

  5. v0.76.0-rc9v0.76.0-rc9Sep 2, 2026pre-release6.6K downloads

    ## What's Changed * Fix crates.io publishing for new workspace crates, and make dev builds identify themselves by @michaelneale in https://github.com/Mesh-LLM/mesh-llm/pull/1223 * Split Skippy by functional boundary by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1194 * feat: pass reasoning effort through Skippy templates by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1211 * chore: update pinned llama.cpp revision by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1216 * fix: IPC idle timeout for external-relay plugins by @MahdiHedhli in https://github.com/Mesh-LLM/mesh-llm/pull/1191 * Publish Rust crate API docs with the website by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1230 * feat(mesh): expose model-bound tokenizer capability by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1227 * fix: repair the Fly console image build by @michaelneale in https://github.com/Mesh-LLM/mesh-llm/pull/1241 * chore(llama): regenerate patch queue for latest upstream by @i386 in https://github.com/Mesh-LLM/mesh-llm/pull/1232 * feat(catalog): add Muse Glimmer 30B to the model catalog by @michaelneale in https://github.com/Mesh-LLM/mesh-llm/pull/1240 * feat(bench): run SW

Commits per week

last 52 weeks
2820Week of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 40 commitsWeek of 2026-02-15: 79 commitsWeek of 2026-02-22: 97 commitsWeek of 2026-03-01: 46 commitsWeek of 2026-03-08: 55 commitsWeek of 2026-03-15: 64 commitsWeek of 2026-03-22: 145 commitsWeek of 2026-03-29: 282 commitsWeek of 2026-04-05: 188 commitsWeek of 2026-04-12: 164 commitsWeek of 2026-04-19: 54 commitsWeek of 2026-04-26: 30 commitsWeek of 2026-05-03: 61 commitsWeek of 2026-05-10: 51 commitsWeek of 2026-05-17: 49 commitsWeek of 2026-05-24: 95 commitsWeek of 2026-05-31: 40 commitsWeek of 2026-06-07: 32 commitsWeek of 2026-06-14: 27 commitsWeek of 2026-06-21: 19 commitsWeek of 2026-06-28: 30 commitsWeek of 2026-07-05: 7 commitsWeek of 2026-07-12: 39 commitsWeek of 2026-07-19: 32 commitsWeek of 2026-07-26: 41 commitsWeek of 2026-08-02: 35 commitsWeek of 2026-08-09: 79 commitsWeek of 2026-08-16: 44 commitsWeek of 2026-08-23: 55 commitsWeek of 2026-08-30: 95 commitsWeek of 2026-09-06: 82 commitsWeek of 2026-09-13: 69 commitsWeek of 2026-09-20: 47 commitsSep 27, 2025Sep 20, 2026
2.3K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 8 commitsSun 1:00 — 6 commitsSun 2:00 — 18 commitsSun 3:00 — 2 commitsSun 4:00 — 6 commitsSun 5:00 — 6 commitsSun 6:00 — 11 commitsSun 7:00 — 13 commitsSun 8:00 — 18 commitsSun 9:00 — 18 commitsSun 10:00 — 16 commitsSun 11:00 — 15 commitsSun 12:00 — 12 commitsSun 13:00 — 11 commitsSun 14:00 — 8 commitsSun 15:00 — 12 commitsSun 16:00 — 11 commitsSun 17:00 — 10 commitsSun 18:00 — 6 commitsSun 19:00 — 9 commitsSun 20:00 — 8 commitsSun 21:00 — 11 commitsSun 22:00 — 8 commitsSun 23:00 — 11 commitsMon 0:00 — 5 commitsMon 1:00 — 7 commitsMon 2:00 — 5 commitsMon 3:00 — 6 commitsMon 4:00 — 2 commitsMon 5:00 — 10 commitsMon 6:00 — 14 commitsMon 7:00 — 7 commitsMon 8:00 — 15 commitsMon 9:00 — 31 commitsMon 10:00 — 27 commitsMon 11:00 — 22 commitsMon 12:00 — 19 commitsMon 13:00 — 8 commitsMon 14:00 — 23 commitsMon 15:00 — 29 commitsMon 16:00 — 26 commitsMon 17:00 — 15 commitsMon 18:00 — 13 commitsMon 19:00 — 16 commitsMon 20:00 — 9 commitsMon 21:00 — 9 commitsMon 22:00 — 9 commitsMon 23:00 — 4 commitsTue 0:00 — 8 commitsTue 1:00 — 3 commitsTue 2:00 — 8 commitsTue 3:00 — 6 commitsTue 4:00 — 2 commitsTue 5:00 — 5 commitsTue 6:00 — 13 commitsTue 7:00 — 12 commitsTue 8:00 — 10 commitsTue 9:00 — 12 commitsTue 10:00 — 18 commitsTue 11:00 — 14 commitsTue 12:00 — 23 commitsTue 13:00 — 15 commitsTue 14:00 — 28 commitsTue 15:00 — 30 commitsTue 16:00 — 21 commitsTue 17:00 — 13 commitsTue 18:00 — 22 commitsTue 19:00 — 25 commitsTue 20:00 — 13 commitsTue 21:00 — 20 commitsTue 22:00 — 7 commitsTue 23:00 — 1 commitsWed 0:00 — 2 commitsWed 1:00 — 3 commitsWed 2:00 — 9 commitsWed 3:00 — 1 commitsWed 4:00 — 3 commitsWed 5:00 — 4 commitsWed 6:00 — 4 commitsWed 7:00 — 12 commitsWed 8:00 — 12 commitsWed 9:00 — 27 commitsWed 10:00 — 16 commitsWed 11:00 — 28 commitsWed 12:00 — 22 commitsWed 13:00 — 26 commitsWed 14:00 — 25 commitsWed 15:00 — 22 commitsWed 16:00 — 13 commitsWed 17:00 — 17 commitsWed 18:00 — 19 commitsWed 19:00 — 18 commitsWed 20:00 — 16 commitsWed 21:00 — 17 commitsWed 22:00 — 14 commitsWed 23:00 — 9 commitsThu 0:00 — 6 commitsThu 1:00 — 8 commitsThu 2:00 — 4 commitsThu 3:00 — 5 commitsThu 4:00 — 6 commitsThu 5:00 — 5 commitsThu 6:00 — 10 commitsThu 7:00 — 15 commitsThu 8:00 — 11 commitsThu 9:00 — 21 commitsThu 10:00 — 23 commitsThu 11:00 — 23 commitsThu 12:00 — 21 commitsThu 13:00 — 29 commitsThu 14:00 — 23 commitsThu 15:00 — 31 commitsThu 16:00 — 33 commitsThu 17:00 — 34 commitsThu 18:00 — 22 commitsThu 19:00 — 14 commitsThu 20:00 — 18 commitsThu 21:00 — 9 commitsThu 22:00 — 10 commitsThu 23:00 — 7 commitsFri 0:00 — 6 commitsFri 1:00 — 8 commitsFri 2:00 — 3 commitsFri 3:00 — 6 commitsFri 4:00 — 14 commitsFri 5:00 — 2 commitsFri 6:00 — 10 commitsFri 7:00 — 32 commitsFri 8:00 — 16 commitsFri 9:00 — 9 commitsFri 10:00 — 16 commitsFri 11:00 — 27 commitsFri 12:00 — 31 commitsFri 13:00 — 27 commitsFri 14:00 — 36 commitsFri 15:00 — 17 commitsFri 16:00 — 26 commitsFri 17:00 — 39 commitsFri 18:00 — 32 commitsFri 19:00 — 18 commitsFri 20:00 — 8 commitsFri 21:00 — 13 commitsFri 22:00 — 19 commitsFri 23:00 — 16 commitsSat 0:00 — 8 commitsSat 1:00 — 6 commitsSat 2:00 — 6 commitsSat 3:00 — 3 commitsSat 4:00 — 5 commitsSat 5:00 — 4 commitsSat 6:00 — 18 commitsSat 7:00 — 20 commitsSat 8:00 — 20 commitsSat 9:00 — 21 commitsSat 10:00 — 19 commitsSat 11:00 — 23 commitsSat 12:00 — 19 commitsSat 13:00 — 16 commitsSat 14:00 — 18 commitsSat 15:00 — 20 commitsSat 16:00 — 14 commitsSat 17:00 — 25 commitsSat 18:00 — 17 commitsSat 19:00 — 18 commitsSat 20:00 — 12 commitsSat 21:00 — 14 commitsSat 22:00 — 12 commitsSat 23:00 — 9 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Jul 12, 2026daily#11+3
  • openclaw/openclaw

    The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

    391.3K stars · TypeScript

  • obra/superpowers

    An agentic skills framework & software development methodology that works.

    295.2K stars · Shell

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    272.8K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    251.2K stars · Python

  • n8n-io/n8n

    Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.

    206.7K stars · TypeScript

  • tensorflow/tensorflow

    An Open Source Machine Learning Framework for Everyone

    200.7K stars · C++