youssofal/MTPLXPublic

3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.

AI summary: A native Mac app and CLI tool for running local language models significantly faster utilizing multi-token prediction.

Stars
2.2K
+52 today
Forks
163
Watchers
18
Open issues
32
Open PRs
45
Contributors
~28
Commits
936
Branches
3

PythonApache-2.0Created May 2, 2026Last push 4d agoLatest release v2.10.2+211 stars this week+327 this month

Star history

since May 3, 2026
01K2KMay 2026Jun 2026Jul 2026Sep 2026
2.2K stars as of Sep 10, 2026, tracked back to May 3, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
SepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 0 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 0 commits2026-02-17: 0 commits2026-02-18: 0 commits2026-02-19: 0 commits2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 0 commits2026-02-24: 0 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 0 commits2026-03-04: 0 commits2026-03-05: 0 commits2026-03-06: 0 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 0 commits2026-03-10: 0 commits2026-03-11: 0 commits2026-03-12: 0 commits2026-03-13: 0 commits2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 0 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 0 commits2026-03-24: 0 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 0 commits2026-04-08: 0 commits2026-04-09: 0 commits2026-04-10: 0 commits2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 0 commits2026-04-15: 0 commits2026-04-16: 0 commits2026-04-17: 0 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 0 commits2026-04-21: 0 commits2026-04-22: 0 commits2026-04-23: 0 commits2026-04-24: 0 commits2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 0 commits2026-04-28: 0 commits2026-04-29: 0 commits2026-04-30: 0 commits2026-05-01: 0 commits2026-05-02: 35 commits2026-05-03: 3 commits2026-05-04: 23 commits2026-05-05: 13 commits2026-05-06: 6 commits2026-05-07: 26 commits2026-05-08: 5 commits2026-05-09: 4 commits2026-05-10: 14 commits2026-05-11: 6 commits2026-05-12: 5 commits2026-05-13: 5 commits2026-05-14: 5 commits2026-05-15: 6 commits2026-05-16: 0 commits2026-05-17: 5 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 0 commits2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 0 commits2026-05-27: 0 commits2026-05-28: 0 commits2026-05-29: 0 commits2026-05-30: 0 commits2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 0 commits2026-06-05: 0 commits2026-06-06: 0 commits2026-06-07: 0 commits2026-06-08: 0 commits2026-06-09: 0 commits2026-06-10: 0 commits2026-06-11: 20 commits2026-06-12: 3 commits2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 1 commit2026-06-18: 1 commit2026-06-19: 1 commit2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 0 commits2026-06-23: 0 commits2026-06-24: 8 commits2026-06-25: 3 commits2026-06-26: 0 commits2026-06-27: 1 commit2026-06-28: 0 commits2026-06-29: 0 commits2026-06-30: 0 commits2026-07-01: 0 commits2026-07-02: 10 commits2026-07-03: 0 commits2026-07-04: 4 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 3 commits2026-07-08: 2 commits2026-07-09: 5 commits2026-07-10: 1 commit2026-07-11: 0 commits2026-07-12: 0 commits2026-07-13: 0 commits2026-07-14: 0 commits2026-07-15: 0 commits2026-07-16: 0 commits2026-07-17: 5 commits2026-07-18: 7 commits2026-07-19: 1 commit2026-07-20: 3 commits2026-07-21: 3 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 1 commit2026-07-25: 0 commits2026-07-26: 12 commits2026-07-27: 0 commits2026-07-28: 4 commits2026-07-29: 5 commits2026-07-30: 4 commits2026-07-31: 54 commits2026-08-01: 26 commits2026-08-02: 35 commits2026-08-03: 10 commits2026-08-04: 3 commits2026-08-05: 0 commits2026-08-06: 3 commits2026-08-07: 4 commits2026-08-08: 16 commits2026-08-09: 39 commits2026-08-10: 10 commits2026-08-11: 7 commits2026-08-12: 1 commit2026-08-13: 0 commits2026-08-14: 26 commits2026-08-15: 31 commits2026-08-16: 27 commits2026-08-17: 19 commits2026-08-18: 12 commits2026-08-19: 0 commits2026-08-20: 5 commits2026-08-21: 6 commits2026-08-22: 22 commits2026-08-23: 0 commits2026-08-24: 12 commits2026-08-25: 23 commits2026-08-26: 73 commits2026-08-27: 41 commits2026-08-28: 41 commits2026-08-29: 36 commits2026-08-30: 6 commits2026-08-31: 14 commits2026-09-01: 2 commits2026-09-02: 0 commits2026-09-03: 0 commits2026-09-04: 0 commits2026-09-05: 0 commits
873 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Rising fast

    +211 stars this week

  • Very active

    873 commits in 52 weeks

  • Well documented

    High community health score

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    9 trending appearances

What MTPLX does

MTPLX is a highly optimized runtime specifically engineered for Apple Silicon hardware to execute local language models. It leverages the multi-token prediction (MTP) heads built into modern models like Qwen to dramatically accelerate generation speeds without requiring secondary draft models. The system drafts several tokens ahead and verifies them in a single batched forward pass, drastically reducing memory overhead. By utilizing exact rejection sampling and residual correction, the output distribution remains mathematically identical to standard sampling. This results in measured speedups of up to 2.24x on M5 Max processors.

This tool is aimed at AI enthusiasts, developers, and researchers utilizing Apple Silicon. It is highly beneficial for those wanting to run powerful models locally without suffering from severe latency or RAM bottlenecks.

  • Multi token prediction: Utilizes built-in MTP heads to draft and verify multiple tokens simultaneously.
  • Apple Silicon optimized: Specifically tuned to extract maximum performance from M-series Mac processors.
  • Exact rejection sampling: Maintains the exact output distribution of the model without employing greedy shortcuts.
  • Single model efficiency: Avoids the heavy RAM usage associated with traditional dual-model speculative decoding.
  • Native interface: Provides both a sleek native Mac application and a flexible command line interface.

Where teams use it

Fast local inference

Developers run large language models directly on their MacBooks with speeds rivaling cloud endpoints.

Memory constrained environments

Users run advanced AI capabilities on 16GB machines without crashing due to excessive RAM overhead.

Offline AI development

Engineers build and test LLM integrations during travel or in secure, air-gapped environments.

High speed code generation

Programmers integrate the fast local endpoint into their IDEs for near-instant coding assistance.

README

main branch
MTPLX

Run local LLMs on Apple Silicon, around twice as fast.

PyPI CI Python macOS Apple Silicon License

MTPLX is a native Mac app and a command line for running local language models with multi-token prediction. Modern models like Qwen 3.5/3.6/3.8 ship with built-in MTP heads. Almost no runtime uses them. MTPLX does: the model drafts several tokens ahead of itself, verifies each drafted block in a single batched forward pass, and commits tokens through exact rejection sampling with residual correction. Same model, same output distribution, measured 1.6x faster on a 16 GB M4 Mac mini and 2.24x on an M5 Max.

There is no second draft model eating your RAM, and no greedy shortcut that quietly changes what the model would have said at real sampling settings. The acceptance math is the Leviathan and Chen rejection sampling theorem with residual correction, so temperature=0.6, top_p=0.95 behaves exactly like normal decoding, just faster.

Get it

The Mac app is the easiest way in. Download the DMG at mtplx.com, drag it to Applications, and the app takes care of everything else: it checks your hardware, recommends a model that actually fits your memory, downloads it, sets up its own Python engine (no Homebrew needed), installs fan control, puts mtplx on your PATH, and then measures your machine to pick the fastest decoding depth.

Recommended for coding: Qwen 3.8 27B Optimized Speed is a 4-bit dynamic quant with great coding speeds and good quality. Its two siblings sit right under it in the app and CLI: Bare Speed (quickest burst chat speeds, lower quality and slower on long coding tasks) and Optimized Quality (8-bit dynamic quant, good coding speeds and perfect quality). Qwen 3.6 Optimized Speed V2 remains available directly below them.

The CLI on its own:

brew install youssofal/mtplx/mtplx
mtplx start

or python3 -m pip install mtplx if you prefer pip. All releases are listed at mtplx.com/releases.

Requirements: Apple Silicon (M1 or newer), macOS 14+. 16 GB of memory runs the 4B and 9B models comfortably. Qwen 3.8 Optimized Speed is recommended on Macs with 32 GB or more; on M1 and M2 the app and CLI pick its FP16 build (same weights, native precision for those chips) automatically. Both check your Mac before recommending anything.

The app

MTPLX dashboard with live decode gauge

The dashboard shows what your model is doing while it does it: live tokens per second, acceptance rate by draft depth, the verify waterfall, cache state, and system pressure. When you start a chat, code an agent against the local server, or run a benchmark, the numbers are right there.

Chat streaming with live speed badge

Chat is native, streams with thinking cards, takes file attachments, and can search the web. One click launches OpenCode, Pi, Hermes, Open WebUI, or anything else that speaks the OpenAI or Anthropic API against your local server. There is also a built-in AIME benchmark runner with fully disclosed, coaching-free prompts, so you can score a model yourself instead of trusting a chart.

Auto-tune

The right draft depth depends on your specific Mac: chip, memory bandwidth, thermals. During onboarding (and any time after), MTPLX runs the real model on your machine at each depth, with fans pinned for clean timing, and keeps autoregressive decoding as the baseline. If an MTP depth beats it, that depth is saved. If nothing beats the baseline, nothing is saved and the app says so. From the terminal it is one command:

mtplx tune --model <model-or-path> --retune

On a 16 GB M4 Mac mini, tuning the 9B model lands on depth 1: 14.4 tok/s baseline becomes 23.0 tok/s.

Forge: make your own MTP models

Forge verifying a freshly built MTP model

Forge takes a Hugging Face repo and turns it into an MTPLX-ready MTP model: convert to MLX, train the MTP adapter, verify that the result is actually faster and still exact, and publish back to the Hub if you want to share it. The honest part matters: Forge measures before and after on your hardware and shows you the verdict ("Depth 1 is fastest: 227.1 to 296.1, 1.30x") rather than assuming the adapter helped. Available in the app and as mtplx forge subcommands.

MTPLX does not support attaching a separately supplied MTP sidecar to an arbitrary MLX trunk. Matching architecture fields, tensor shapes, or provenance labels cannot prove that the head was trained against those exact trunk weights. Use a complete model that already includes its matching MTP weights, or use Forge to build and verify an artifact from its original source checkpoint.

The official catalog lives on Hugging Face under Youssofal: Qwen 3.8 27B (Bare Speed, Optimized Speed, Optimized Quality, each with an FP16 build for M1 and M2), Qwen 3.6 (27B, 35B MoE) in speed and quality builds (the 35B MoE adds a balance build), Qwen 3.5 (4B, 9B), plus Gemma 4. The app and the CLI recommend from these based on your hardware.

The server

mtplx start (or the app's play button) serves an OpenAI-compatible API on 127.0.0.1:8000: /v1/chat/completions, /v1/completions, /v1/models, the optional /v1/embeddings and /v1/rerank (see below), plus an Anthropic-compatible /v1/messages with streaming, tool calls in both styles, /health, and /metrics. Claude Code, Cline, Continue, Open WebUI, curl, the openai and anthropic Python clients: if it speaks the API, it works. The app and CLI share one server, so mtplx start attaches to the app's running model instead of loading a second copy.

curl http://127.0.0.1:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"mtplx","messages":[{"role":"user","content":"hi"}],"stream":true}'

Sessions survive: a warm-prefix session bank keeps multi-turn chats fast, and a default-on SSD session cache restores sessions near-instantly across restarts (disable with --ssd-session-cache off).

Embeddings and reranking

The same daemon can serve retrieval models, so a RAG or agent-memory setup does not need a second inference server beside MTPLX. Point it at any MLX embedding or reranker model — Hugging Face id or local path, optionally with a REF=served-id alias:

mtplx serve \
  --embedding-model mlx-community/Qwen3-Embedding-8B-4bit-DWQ \
  --reranker-model vserifsaglam/Qwen3-Reranker-4B-4bit-MLX
curl http://127.0.0.1:8000/v1/embeddings \
  -H 'Content-Type: application/json' \
  -d '{"model":"Qwen3-Embedding-8B-4bit-DWQ","input":["hello","world"]}'

curl http://127.0.0.1:8000/v1/rerank \
  -H 'Content-Type: application/json' \
  -d '{"query":"where is the cache?","documents":["the cache lives in ~/.mtplx","unrelated text"]}'

Both flags repeat, so several models can be served at once and picked per request via "model". Listing the same reference as both an embedder and a reranker loads one copy of the weights and serves both roles from it. Retrieval models load on first request and are capped by --retrieval-max-resident (default 2), which unloads the least recently used one beyond the cap — an unused endpoint costs nothing. /v1/models stays chat-only by default so chat clients that enumerate models never offer an embedder as a conversation target; list retrieval models with ?capability=embedding or ?capability=rerank (every entry carries its capability), and a chat completion that requests a retrieval id gets a clear 400 rather than a silent answer from the chat model.

These models do not go through the MTP path, and that is deliberate: multi-token prediction makes next-token decoding cheaper, which means nothing for a model that returns a vector instead of a token stream. Configure them in the app under Settings → Retrieval endpoints, or persist them in ~/.mtplx/config.toml as embedding_models and reranker_models. With nothing configured the endpoints answer 404 and chat behaves exactly as before. One safety gate: checkpoints that bundle their own Python inference code (the jina embedding/reranker MLX releases do) are refused with a 403 until you opt in with --retrieval-trust-remote-code (or retrieval_trust_remote_code = true in the config file) — a model download never gains code execution just by being pointed at.

Sampler controls cover temperature, top_p, top_k, and the OpenAI penalty pair presence_penalty / frequency_penalty — per request, as server defaults (--default-presence-penalty / --default-frequency-penalty on start/serve/quickstart), or live via mtplx settings set and the app's Presence Penalty dial. Penalties default to 0, which is an exact no-op that preserves MTP exactness. Qwen's guidance: leave them at 0 for coding and agent work; ~0.5–1.5 presence penalty helps creative writing or when a model loops on itself.

Concurrent scheduler modes, ownership guarantees, and backend-specific implementations are documented in Concurrency modes.

CLI quick reference

mtplx start                # interactive: pick model, mode, surface, then chat
mtplx serve --port 8000    # API server only
mtplx stop                 # stop the running server cleanly
mtplx pull <hf-repo>       # download a model safely
mtplx models               # what is cached, sizes, validation
mtplx inspect <model>      # compatibility report before anything runs
mtplx tune --retune        # measure AR vs D1/D2/D3 on your Mac
mtplx forge --help         # build, verify, and publish MTP models (probe/build/publish/verify subcommands)
mtplx bench aime --quick   # run the AIME benchmark from the terminal
mtplx doctor               # install and integration health
mtplx max --install        # fan control (one sudo prompt, crash-safe)
mtplx settings get/set     # read or change live server settings

Every command takes --help, and most inspection/diagnostic commands take --json. The CLI works without MLX installed for everything that does not need a model, so doctor and inspect run on any machine.

Modes

Mode What it does When
Turbo NAX verify kernels + compiled verify; the default for the quantized 27B and 9B flagship models Picked automatically for those models
Sustained Default for all other models. Long-context MTP path with chunked prefill and request-sized KV Everyday use, big files, 16K-200K prompts
Sustained Max Sustained with fans pinned at 100% Long work where you want maximum cooling
Burst Legacy short-context benchmark lane, loud Short prompts and benchmarks only

Fan-backed modes restore your fans to automatic if MTPLX dies for any reason, including kill -9 and closing the terminal. A detached watchdog handles it; this is verified on hardware, not assumed.

Compatibility, honestly

mtplx inspect classifies models before anything runs: verified, family-compatible but unverified, architecture-compatible but unverified, AR-only, incompatible architecture, or no MTP heads at all. Unverified models load with an explicit unverified label. There are no silent fallbacks: if MTPLX cannot run a model correctly, it tells you instead of running it badly.

Laguna-S-2.1 oQ4e is supported through its exact MLX architecture in target-only AR mode:

mtplx start cli \
  --model mlx-community/Laguna-S-2.1-oQ4e \
  --download \
  --no-mtp

MTPLX pins that model to revision 8e3f5cad513746264940c1c4195de48d7ea345a5 and verifies the 13-shard layout, tokenizer, generation config, special tokens map, and Poolside chat template before admitting it. The checkpoint has no native MTP head, so an MTP launch is rejected before weights load instead of falling back during execution. The weights occupy 59.72 GiB, a 64.13 GB snapshot on disk. The launch preflight requires about 85 GiB of unified memory (weights, runtime headroom, and a 16 GiB system reserve) — in practice a 96 GB Mac; 128 GB is comfortable. MTPLX defaults Laguna to a 32,768-token context and response cap, and checks larger explicit server contexts against the active Metal memory cap.

What MTPLX is not

  • Not an external-drafter system. The drafter is the target model's own MTP heads.
  • Not a greedy-argmax trick. Acceptance is exact rejection sampling, correct at any temperature.
  • Not a CUDA project. MTPLX is MLX-native and Apple Silicon first. For Linux, use vLLM.

History

MTPLX was the first runtime on Apple Silicon to run a model's own MTP heads with mathematically exact speculative sampling — 27 April 2026, before llama.cpp had MTP at all, and months before it reached the hybrid GDN family. The dated record, with a public receipt for every claim, is in HISTORY.md and at mtplx.com/history.

License and credit

Apache-2.0: use it, modify it, ship it commercially. Keep the license and the NOTICE file if you redistribute.

Attribution is required. If you ship a product, app, or service that includes or is built on MTPLX, it has to say so inside the product itself, somewhere a user can see it (About screen, credits, settings, shipped docs, or a CLI startup banner):

Powered by MTPLX https://github.com/youssofal/MTPLX

A mention in your repo or on your website does not cover it. The full terms are in NOTICE, which Apache-2.0 section 4(d) carries with every copy.

MTPLX builds on MLX and the Qwen and Gemma model families; the speculative sampling math follows Leviathan and Chen (2023). Fan control via ThermalForge. Model weights remain governed by their upstream licenses.

Built by Youssof Altoukhi. Bug reports and benchmark replications welcome via Issues.

View on GitHub

Recent activity

commits and pull requests

Recent open issues

view all

Releases and announcements

50 total
  1. MTPLX 2.10.2v2.10.2Sep 1, 2026447 downloads

    # MTPLX 2.10.2 Honest memory refusals, a correct and resilient Anthropic bridge for Claude Code, and sharper stop diagnostics. ## Memory admission answers before it wedges (#415) Large prompts used to be admitted optimistically. A request that could not fit prefilled until the allocator hit the wall, the stream died mid-flight, and the failure was logged as a client cancellation. 2.10.2 projects the prefill footprint before admission. Superseded SessionBank entries are cleared proactively, and a request that genuinely cannot fit is answered upfront with a structured HTTP 507 naming the shortfall. When a stream does fail, the wire and the request log now carry an honest error receipt (`stream_error`, `error_kind`) instead of a cancellation entry. ## Anthropic bridge: correct usage, resilient streams Two fixes for Claude Code and any other Anthropic-dialect client. **Usage math (PR #417 by @amichaelblock-lgtm).** Anthropic's `input_tokens` and `cache_read_input_tokens` are disjoint fields; OpenAI's `prompt_tokens` is a cumulative total. The bridge copied the cumulative total into `input_tokens` while also reporting `cache_read_input_tokens`, double-counting the cached prefix o

  2. MTPLX 2.10.1v2.10.1Aug 30, 20262.9K downloads

    # MTPLX 2.10.1 Patch release: faster long-prompt processing on M4 and M5 Macs, working image input for Flash-Next, and fixes for 96 GB and M2/M3 machines. ## Prefill Flash-Next gains a sparse prefill lane, adapting the fused Metal kernels from PR #397 by @maceip. The model already scores which blocks of the context matter for each new token; the lane now feeds those scores straight into a block-sparse FlashAttention kernel instead of building full attention masks. Measured on an M5 Max with 128 GB against 2.10.0 on the same machine: - 98k-token prompt: processing time drops 35 percent (175.7 s to 114.5 s), peak memory drops from 91.4 to 83.0 GB. - 131k-token prompt: 810 tok/s. - 262,144-token cold prompt: completes in 355 s at 87.4 GB peak. On 2.10.0 the same request climbed to 119 GB and produced nothing (#393, reported by @blackjose007-stack). The lane turns on automatically where its kernels are supported (M4 and M5 generation GPUs) and engages on prompts past 32k tokens; other machines keep the dense path. `MTPLX_QSA_PREFILL=0` turns it off. The kernels also handle YaRN rope scaling for long-rope configurations. Memory admission understands the lane: serve resolves the m

  3. MTPLX 2.10.0v2.10.0Aug 29, 20262K downloads

    # MTPLX 2.10.0 Measured on an M5 Max against 2.9.2, stock settings: | | 2.9.2 | 2.10.0 | |---|---:|---:| | 27B decode, 3k chat answer | 55.9 tok/s | 64.3 tok/s (+15%) | | 27B decode, 88k context | 23.6 tok/s | 30.4 tok/s (+29%) | | 27B decode, 147k context | 12.0 tok/s | 18.4 tok/s (+54%) | | 27B prefill, 88k context | 379 tok/s | 535 tok/s (+41%) | | Decode cost with q8 KV quantization | crash, or -50% | -4% | | Decode cost with q4 KV quantization | crash, or -50% | -19% | | Rewriting a file the model just wrote | 73.8 tok/s | 87.6 tok/s (+19%) | | Allocator growth in one long answer | 8.6 GB | 0.6 GB (-93%) | | Multi-file agent task, wall clock | 150.2 s | 44.2 s (-71%) | | Agent first token, mid-session | 1.8 to 2.2 s | 0.11 s (-94%) | | 48 GB Mac, long coding session | 3.1 to 4.6 tok/s in swap (#305) | 33 tok/s at 42k (7 to 10x) | Every row is stock 2.9.2 against stock 2.10.0 on the same Mac, so rows that ride a raised dependency floor or a changed default show that full stock-to-stock gain; the body section for each row says exactly what was measured and how. The 48 GB row is the pinned simulated seat from the release gate. ## Qwen 3.8 Flash-Next Qwen's 125B MoE preview i

  4. MTPLX 2.9.2v2.9.2Aug 25, 20264.9K downloads

    # MTPLX 2.9.2 MTPLX stops rewriting agent transcripts, greedy decoding gets faster below 12k context, and the model forge gets a correctness fix that rescues packs whose draft acceptance had collapsed. ## Your transcript is yours (#282) - The serving endpoints are passthrough by default. MTPLX no longer compacts tool results, trims file reads, or injects steering text into agent transcripts unless you explicitly turn a rewrite feature on. `MTPLX_AGENT_REWRITES` is the master switch, and each individual feature only arms when you set its own environment variable. - The macOS app stopped exporting the legacy compaction settings when it launches coding agents, so app-launched Pi and OpenCode sessions get the same clean passthrough as the CLI. - Managed client configs respect your edits. `mtplx start` and the app only update files they wrote themselves, and never overwrite a config you have customized. - The request log records exactly what was and was not rewritten on every request, so you can verify the passthrough yourself. ## Faster greedy decode below 12k context Chained greedy drafting is now on by default for temperature 0 requests with prompts under 12,288 tokens (#313, #3

  5. MTPLX 2.9.1v2.9.1Aug 22, 20263.8K downloads

    # MTPLX 2.9.1 Agent coding sessions run to completion: long-context crash fixes, no hidden output caps, reasoning preserved across turns, and a built-in flight recorder for diagnosing any session. ## Engine - **Fixed: agent sessions could truncate and crash near 19,000 tokens** (#310). The paged KV cache derived its capacity from a stompable claim instead of the pages it had actually allocated. Long coding sessions now run to the model's full advertised context. - **Fixed: shutdown segfault** (#303). The daemon parks its model-owner thread and clears MLX streams at exit, so quit and restart are clean. - **Turbo profile truth.** 2.9.0 shipped one turbo fast-path flag that was runtime-dead, so turbo did not apply its full intended configuration. The fast-path environment is now a single shared block, `/health` reports exactly what the profile set, and a per-lane kernel selfcheck runs at startup. If you benchmarked turbo on 2.9.0, re-run it. - **Multi-turn cache reuse holds at scale.** All encode paths now share one tokenization policy, so warm agent turns no longer hit cache walls at assistant reasoning boundaries; tool-call turns bank their just-generated output directly from liv

Code frequency

additions and deletions
+168.9K-168.9KWeek of 2026-04-26: +46,375 linesWeek of 2026-04-26: -2,812 linesWeek of 2026-05-03: +70,659 linesWeek of 2026-05-03: -3,798 linesWeek of 2026-05-10: +12,905 linesWeek of 2026-05-10: -1,010 linesWeek of 2026-05-17: +4,222 linesWeek of 2026-05-17: -867 linesWeek of 2026-05-24: +0 linesWeek of 2026-05-24: -0 linesWeek of 2026-05-31: +0 linesWeek of 2026-05-31: -0 linesWeek of 2026-06-07: +168,902 linesWeek of 2026-06-07: -6,507 linesWeek of 2026-06-14: +228 linesWeek of 2026-06-14: -41 linesWeek of 2026-06-21: +594 linesWeek of 2026-06-21: -146 linesWeek of 2026-06-28: +17,942 linesWeek of 2026-06-28: -2,580 linesWeek of 2026-07-05: +5,840 linesWeek of 2026-07-05: -257 linesWeek of 2026-07-12: +11,744 linesWeek of 2026-07-12: -855 linesWeek of 2026-07-19: +6,384 linesWeek of 2026-07-19: -170 linesWeek of 2026-07-26: +69,616 linesWeek of 2026-07-26: -2,815 linesWeek of 2026-08-02: +36,505 linesWeek of 2026-08-02: -1,259 linesWeek of 2026-08-09: +32,619 linesWeek of 2026-08-09: -4,102 linesWeek of 2026-08-16: +36,572 linesWeek of 2026-08-16: -3,596 linesWeek of 2026-08-23: +49,923 linesWeek of 2026-08-23: -3,153 linesWeek of 2026-08-30: +3,319 linesWeek of 2026-08-30: -217 linesApr 26, 2026Aug 30, 2026
+574.3K lines added, -34.2K removed over the last year.

Commits per week

last 52 weeks
2260Week of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 0 commitsWeek of 2026-02-15: 0 commitsWeek of 2026-02-22: 0 commitsWeek of 2026-03-01: 0 commitsWeek of 2026-03-08: 0 commitsWeek of 2026-03-15: 0 commitsWeek of 2026-03-22: 0 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 0 commitsWeek of 2026-04-12: 0 commitsWeek of 2026-04-19: 0 commitsWeek of 2026-04-26: 35 commitsWeek of 2026-05-03: 80 commitsWeek of 2026-05-10: 41 commitsWeek of 2026-05-17: 5 commitsWeek of 2026-05-24: 0 commitsWeek of 2026-05-31: 0 commitsWeek of 2026-06-07: 23 commitsWeek of 2026-06-14: 3 commitsWeek of 2026-06-21: 12 commitsWeek of 2026-06-28: 14 commitsWeek of 2026-07-05: 11 commitsWeek of 2026-07-12: 12 commitsWeek of 2026-07-19: 8 commitsWeek of 2026-07-26: 105 commitsWeek of 2026-08-02: 71 commitsWeek of 2026-08-09: 114 commitsWeek of 2026-08-16: 91 commitsWeek of 2026-08-23: 226 commitsWeek of 2026-08-30: 22 commitsSep 7, 2025Aug 30, 2026
873 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 11 commitsSun 1:00 — 9 commitsSun 2:00 — 9 commitsSun 3:00 — 6 commitsSun 4:00 — 3 commitsSun 5:00 — 4 commitsSun 6:00 — 8 commitsSun 7:00 — 5 commitsSun 8:00 — 6 commitsSun 9:00 — 0 commitsSun 10:00 — 2 commitsSun 11:00 — 1 commitsSun 12:00 — 2 commitsSun 13:00 — 7 commitsSun 14:00 — 4 commitsSun 15:00 — 6 commitsSun 16:00 — 10 commitsSun 17:00 — 13 commitsSun 18:00 — 3 commitsSun 19:00 — 6 commitsSun 20:00 — 6 commitsSun 21:00 — 5 commitsSun 22:00 — 4 commitsSun 23:00 — 12 commitsMon 0:00 — 9 commitsMon 1:00 — 2 commitsMon 2:00 — 1 commitsMon 3:00 — 6 commitsMon 4:00 — 3 commitsMon 5:00 — 4 commitsMon 6:00 — 9 commitsMon 7:00 — 0 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 0 commitsMon 11:00 — 0 commitsMon 12:00 — 2 commitsMon 13:00 — 1 commitsMon 14:00 — 12 commitsMon 15:00 — 0 commitsMon 16:00 — 3 commitsMon 17:00 — 6 commitsMon 18:00 — 3 commitsMon 19:00 — 11 commitsMon 20:00 — 4 commitsMon 21:00 — 9 commitsMon 22:00 — 3 commitsMon 23:00 — 10 commitsTue 0:00 — 10 commitsTue 1:00 — 9 commitsTue 2:00 — 6 commitsTue 3:00 — 4 commitsTue 4:00 — 5 commitsTue 5:00 — 1 commitsTue 6:00 — 4 commitsTue 7:00 — 4 commitsTue 8:00 — 2 commitsTue 9:00 — 2 commitsTue 10:00 — 2 commitsTue 11:00 — 4 commitsTue 12:00 — 5 commitsTue 13:00 — 0 commitsTue 14:00 — 1 commitsTue 15:00 — 0 commitsTue 16:00 — 2 commitsTue 17:00 — 2 commitsTue 18:00 — 4 commitsTue 19:00 — 3 commitsTue 20:00 — 0 commitsTue 21:00 — 3 commitsTue 22:00 — 1 commitsTue 23:00 — 1 commitsWed 0:00 — 1 commitsWed 1:00 — 0 commitsWed 2:00 — 2 commitsWed 3:00 — 3 commitsWed 4:00 — 4 commitsWed 5:00 — 13 commitsWed 6:00 — 8 commitsWed 7:00 — 1 commitsWed 8:00 — 4 commitsWed 9:00 — 0 commitsWed 10:00 — 0 commitsWed 11:00 — 0 commitsWed 12:00 — 3 commitsWed 13:00 — 5 commitsWed 14:00 — 0 commitsWed 15:00 — 8 commitsWed 16:00 — 11 commitsWed 17:00 — 5 commitsWed 18:00 — 4 commitsWed 19:00 — 7 commitsWed 20:00 — 8 commitsWed 21:00 — 8 commitsWed 22:00 — 3 commitsWed 23:00 — 3 commitsThu 0:00 — 5 commitsThu 1:00 — 17 commitsThu 2:00 — 6 commitsThu 3:00 — 6 commitsThu 4:00 — 7 commitsThu 5:00 — 12 commitsThu 6:00 — 7 commitsThu 7:00 — 1 commitsThu 8:00 — 1 commitsThu 9:00 — 0 commitsThu 10:00 — 1 commitsThu 11:00 — 0 commitsThu 12:00 — 1 commitsThu 13:00 — 6 commitsThu 14:00 — 4 commitsThu 15:00 — 2 commitsThu 16:00 — 14 commitsThu 17:00 — 5 commitsThu 18:00 — 7 commitsThu 19:00 — 2 commitsThu 20:00 — 0 commitsThu 21:00 — 2 commitsThu 22:00 — 7 commitsThu 23:00 — 11 commitsFri 0:00 — 8 commitsFri 1:00 — 7 commitsFri 2:00 — 4 commitsFri 3:00 — 7 commitsFri 4:00 — 8 commitsFri 5:00 — 12 commitsFri 6:00 — 3 commitsFri 7:00 — 0 commitsFri 8:00 — 1 commitsFri 9:00 — 2 commitsFri 10:00 — 5 commitsFri 11:00 — 5 commitsFri 12:00 — 5 commitsFri 13:00 — 0 commitsFri 14:00 — 3 commitsFri 15:00 — 5 commitsFri 16:00 — 3 commitsFri 17:00 — 9 commitsFri 18:00 — 6 commitsFri 19:00 — 15 commitsFri 20:00 — 8 commitsFri 21:00 — 11 commitsFri 22:00 — 20 commitsFri 23:00 — 6 commitsSat 0:00 — 9 commitsSat 1:00 — 6 commitsSat 2:00 — 16 commitsSat 3:00 — 16 commitsSat 4:00 — 32 commitsSat 5:00 — 11 commitsSat 6:00 — 4 commitsSat 7:00 — 4 commitsSat 8:00 — 1 commitsSat 9:00 — 1 commitsSat 10:00 — 6 commitsSat 11:00 — 0 commitsSat 12:00 — 0 commitsSat 13:00 — 0 commitsSat 14:00 — 1 commitsSat 15:00 — 1 commitsSat 16:00 — 9 commitsSat 17:00 — 13 commitsSat 18:00 — 12 commitsSat 19:00 — 4 commitsSat 20:00 — 4 commitsSat 21:00 — 14 commitsSat 22:00 — 11 commitsSat 23:00 — 7 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits729 (79%)
Community commits192 (21%)

921 commits in total over the last year.

DateListRankStars gained
Sep 10, 2026monthly#11+1,055
Sep 8, 2026monthly#9+1,000
Sep 7, 2026monthly#9+1,000
Sep 6, 2026monthly#11+953
Sep 5, 2026monthly#13+922
Sep 4, 2026monthly#13+870
Sep 3, 2026monthly#12+826
Sep 2, 2026monthly#19+766
Sep 1, 2026monthly#19+746