jundot/omlxPublic

LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar

AI summary: A highly optimized LLM inference server for Apple Silicon, controllable directly from the macOS menu bar.

Stars
18.5K
+39 today
Forks
1.6K
Watchers
97
Open issues
662
Open PRs
147
Contributors
~201
Commits
2.1K
Branches
22

PythonApache-2.0Created Feb 13, 2026Last push todayLatest release v0.5.7+186 stars this week+239 this month

Star history

since Feb 8, 2026
05K10K15KFeb 2026Apr 2026Jun 2026Aug 2026
18.5K stars as of Aug 7, 2026, tracked back to Feb 8, 2026. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-09: 0 commits2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 0 commits2026-02-06: 0 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 0 commits2026-02-10: 0 commits2026-02-11: 0 commits2026-02-12: 0 commits2026-02-13: 10 commits2026-02-14: 9 commits2026-02-15: 8 commits2026-02-16: 8 commits2026-02-17: 3 commits2026-02-18: 1 commit2026-02-19: 1 commit2026-02-20: 0 commits2026-02-21: 0 commits2026-02-22: 3 commits2026-02-23: 3 commits2026-02-24: 11 commits2026-02-25: 11 commits2026-02-26: 4 commits2026-02-27: 20 commits2026-02-28: 20 commits2026-03-01: 14 commits2026-03-02: 26 commits2026-03-03: 21 commits2026-03-04: 15 commits2026-03-05: 14 commits2026-03-06: 13 commits2026-03-07: 7 commits2026-03-08: 16 commits2026-03-09: 6 commits2026-03-10: 21 commits2026-03-11: 37 commits2026-03-12: 18 commits2026-03-13: 18 commits2026-03-14: 21 commits2026-03-15: 34 commits2026-03-16: 14 commits2026-03-17: 28 commits2026-03-18: 12 commits2026-03-19: 11 commits2026-03-20: 2 commits2026-03-21: 36 commits2026-03-22: 28 commits2026-03-23: 5 commits2026-03-24: 17 commits2026-03-25: 10 commits2026-03-26: 14 commits2026-03-27: 12 commits2026-03-28: 16 commits2026-03-29: 43 commits2026-03-30: 7 commits2026-03-31: 7 commits2026-04-01: 5 commits2026-04-02: 3 commits2026-04-03: 25 commits2026-04-04: 8 commits2026-04-05: 21 commits2026-04-06: 12 commits2026-04-07: 17 commits2026-04-08: 6 commits2026-04-09: 3 commits2026-04-10: 21 commits2026-04-11: 5 commits2026-04-12: 2 commits2026-04-13: 3 commits2026-04-14: 32 commits2026-04-15: 13 commits2026-04-16: 7 commits2026-04-17: 6 commits2026-04-18: 0 commits2026-04-19: 3 commits2026-04-20: 2 commits2026-04-21: 6 commits2026-04-22: 21 commits2026-04-23: 7 commits2026-04-24: 8 commits2026-04-25: 1 commit2026-04-26: 0 commits2026-04-27: 6 commits2026-04-28: 16 commits2026-04-29: 0 commits2026-04-30: 4 commits2026-05-01: 1 commit2026-05-02: 0 commits2026-05-03: 2 commits2026-05-04: 25 commits2026-05-05: 4 commits2026-05-06: 19 commits2026-05-07: 4 commits2026-05-08: 0 commits2026-05-09: 15 commits2026-05-10: 3 commits2026-05-11: 18 commits2026-05-12: 37 commits2026-05-13: 24 commits2026-05-14: 7 commits2026-05-15: 8 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 2 commits2026-05-19: 30 commits2026-05-20: 9 commits2026-05-21: 11 commits2026-05-22: 18 commits2026-05-23: 0 commits2026-05-24: 11 commits2026-05-25: 7 commits2026-05-26: 18 commits2026-05-27: 28 commits2026-05-28: 10 commits2026-05-29: 18 commits2026-05-30: 15 commits2026-05-31: 17 commits2026-06-01: 19 commits2026-06-02: 29 commits2026-06-03: 30 commits2026-06-04: 20 commits2026-06-05: 25 commits2026-06-06: 17 commits2026-06-07: 16 commits2026-06-08: 18 commits2026-06-09: 29 commits2026-06-10: 14 commits2026-06-11: 31 commits2026-06-12: 4 commits2026-06-13: 9 commits2026-06-14: 9 commits2026-06-15: 14 commits2026-06-16: 10 commits2026-06-17: 6 commits2026-06-18: 3 commits2026-06-19: 3 commits2026-06-20: 0 commits2026-06-21: 4 commits2026-06-22: 4 commits2026-06-23: 0 commits2026-06-24: 4 commits2026-06-25: 11 commits2026-06-26: 6 commits2026-06-27: 0 commits2026-06-28: 4 commits2026-06-29: 17 commits2026-06-30: 1 commit2026-07-01: 3 commits2026-07-02: 7 commits2026-07-03: 0 commits2026-07-04: 4 commits2026-07-05: 0 commits2026-07-06: 1 commit2026-07-07: 1 commit2026-07-08: 25 commits2026-07-09: 34 commits2026-07-10: 20 commits2026-07-11: 4 commits2026-07-12: 22 commits2026-07-13: 1 commit2026-07-14: 2 commits2026-07-15: 1 commit2026-07-16: 0 commits2026-07-17: 8 commits2026-07-18: 19 commits2026-07-19: 0 commits2026-07-20: 21 commits2026-07-21: 29 commits2026-07-22: 20 commits2026-07-23: 7 commits2026-07-24: 6 commits2026-07-25: 3 commits2026-07-26: 4 commits2026-07-27: 6 commits2026-07-28: 19 commits2026-07-29: 20 commits2026-07-30: 26 commits2026-07-31: 7 commits2026-08-01: 5 commits2026-08-02: 19 commits2026-08-03: 11 commits2026-08-04: 11 commits2026-08-05: 1 commit2026-08-06: 0 commits2026-08-07: 0 commits2026-08-08: 0 commits
2,003 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    18,510 stars

  • Very active

    2,003 commits in 52 weeks

  • Community-driven

    ~201 contributors

  • Outside contributions

    94% of recent commits from the community

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

What omlx does

oMLX is a specialized inference server designed to run Large Language Models efficiently on Apple Silicon (M-series) Macs using the MLX framework. It implements advanced techniques like continuous batching and tiered KV caching, significantly maximizing throughput and memory efficiency compared to naive implementations. The backend exposes an OpenAI-compatible API, allowing it to seamlessly drop into existing workflows and tools. Uniquely, the entire server lifecycle, model downloading, and performance monitoring are managed via a lightweight, native macOS menu bar application, removing the need for complex terminal commands.

Targeted at Mac users with Apple Silicon (M1/M2/M3/M4) chips who want to run powerful local LLMs with high efficiency. It requires macOS and sufficient unified memory to load the desired models.

  • Apple Silicon optimization: Built specifically on Apple's MLX framework to fully utilize unified memory and Neural Engine hardware.
  • OpenAI API compatibility: Serves models via standard endpoints, making it instantly compatible with UI clients and coding agents.
  • Continuous batching: Processes multiple requests simultaneously, drastically improving throughput for concurrent API calls.
  • Native menu bar UI: Provides a simple GUI to start, stop, and monitor models without touching the command line.
  • Tiered KV caching: Intelligently manages context memory to prevent the server from crashing when context windows get large.

Where teams use it

Local AI coding assistance

Developers point their IDE tools like Cursor to the local oMLX server to get zero-latency code completions without paying API fees.

Running private autonomous agents

Researchers execute multi-step agent loops entirely on their Mac, ensuring sensitive data never leaves their machine.

High-throughput local processing

Engineers batch process thousands of text classification tasks locally overnight using the continuous batching capabilities.

Offline model experimentation

Enthusiasts easily download and test the latest open-source models using the intuitive menu bar interface.

Getting started: Download the oMLX.dmg release, drag it to Applications, and select a model from the menu bar.

README

main branch

oMLX

oMLX

LLM inference, optimized for your Mac
Continuous batching and tiered KV caching, managed directly from your menu bar.

Buy Me A Coffee

License Python 3.11-3.13 Apple Silicon

junkim.dot@gmail.com · https://omlx.ai/me

Install · Quickstart · Features · Models · CLI Configuration · Benchmarks · oMLX.ai

English · 中文 · 한국어 · 日本語


oMLX Admin Dashboard

Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar.

oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and reusable across requests, making local LLMs practical for real coding work with tools like Claude Code. That's why I built it.

Install

macOS App

Download the .dmg from Releases, drag to Applications, done. The app includes in-app auto-update, so future upgrades are just one click. The macOS app also installs a lightweight ~/.omlx/bin/omlx CLI shim so terminal commands and Apple Shortcuts can control the app-managed server.

Homebrew

brew tap jundot/omlx https://github.com/jundot/omlx
brew install omlx

# Upgrade to the latest version
brew update && brew upgrade omlx

# Run as a background service (auto-restarts on crash)
omlx start

# Optional: MCP (Model Context Protocol) support
/opt/homebrew/opt/omlx/libexec/bin/pip install mcp

Optional GLM-5.2 / MiniMax M3 native custom kernels currently require a HEAD build:

brew install omlx --HEAD --with-custom-kernel

From Source

git clone https://github.com/jundot/omlx.git
cd omlx
pip install -e .          # Core only
pip install -e ".[mcp]"   # With MCP (Model Context Protocol) support

# GLM-5.2 / MiniMax M3 / Qwen3.5 native custom kernels (strongly recommended
# if you serve those families -- see note below)
OMLX_WITH_CUSTOM_KERNEL=1 pip install -e .

Requires macOS 15.0+ (Sequoia), Python 3.11–3.13, and Apple Silicon (M1/M2/M3/M4).

Note on native custom kernels: a plain pip install -e . does NOT build them, and the affected model families then silently fall back to much slower generic paths -- for GLM-5.2 the fused DSA prefill is roughly 30x faster with the kernels (measured 845 vs ~29 tok/s on an M3 Ultra), and the fallback also uses more memory (#2137). Building them requires the Metal toolchain, which Command Line Tools alone do not provide (xcrun: error: unable to find utility "metal"): install full Xcode, or use the official DMG which ships the kernels precompiled. Homebrew can build them with brew install omlx --HEAD --with-custom-kernel, but that build also needs full Xcode. To verify your install:

python -c "from omlx.custom_kernels import native_kernel_status; print(native_kernel_status())"

Quickstart

macOS App

Launch oMLX from your Applications folder. The Welcome screen guides you through three steps - model directory, server start, and first model download. That's it. To connect OpenClaw, OpenCode, Codex, Hermes Agent, or Copilot, see Integrations.

oMLX Welcome Screen oMLX Menubar

CLI

# Managed background server (macOS app or Homebrew install)
omlx start
omlx stop
omlx restart

# Foreground server attached to this terminal
omlx serve --model-dir ~/models

The server discovers LLMs, VLMs, embedding models, and rerankers from subdirectories automatically. Any OpenAI-compatible client can connect to http://localhost:8000/v1. A built-in chat UI is also available at http://localhost:8000/admin/chat.

Homebrew Service

If you installed via Homebrew, you can run oMLX as a managed background service:

omlx start                    # Start via brew services
omlx stop                     # Stop
omlx restart                  # Restart

brew services start omlx    # Start (auto-restarts on crash)
brew services stop omlx     # Stop
brew services restart omlx  # Restart
brew services info omlx     # Check status

The service runs omlx serve with zero-config defaults (~/.omlx/models, port 8000). omlx start, omlx stop, and omlx restart are the portable lifecycle commands; Homebrew installs delegate them to brew services. To customize, either set environment variables (OMLX_MODEL_DIR, OMLX_PORT, etc.) or run omlx serve --model-dir /your/path once to persist settings to ~/.omlx/settings.json.

Logs are written to two locations:

  • Service log: $(brew --prefix)/var/log/omlx.log (stdout/stderr)
  • Server log: ~/.omlx/logs/server.log (structured application log)

Features

Supports text LLMs, vision-language models (VLM), OCR models, embeddings, and rerankers on Apple Silicon.

Admin Dashboard

Web UI at /admin for real-time monitoring, model management, chat, benchmark, and per-model settings. Supports English, Korean, Japanese, Chinese, French, Russian, Spanish, and Brazilian Portuguese. All CDN dependencies are vendored for fully offline operation.

oMLX Admin Dashboard

Vision-Language Models

Run VLMs with the same continuous batching and tiered KV cache stack as text LLMs. Supports multi-image chat, base64/URL/file image inputs, and tool calling with vision context. OCR models (DeepSeek-OCR, DOTS-OCR, GLM-OCR) are auto-detected with optimized prompts.

Tiered KV Cache (Hot + Cold)

Block-based KV cache management inspired by vLLM, with prefix sharing and Copy-on-Write. The cache operates across two tiers:

  • Hot tier (RAM): Frequently accessed blocks stay in memory for fast access.
  • Cold tier (SSD): When the hot cache fills up, blocks are offloaded to SSD in safetensors format. On the next request with a matching prefix, they're restored from disk instead of recomputed from scratch - even after a server restart.

oMLX Hot & Cold Cache

Continuous Batching

Handles concurrent requests through mlx-lm's BatchGenerator. Max concurrent requests is configurable via CLI or admin panel.

Claude Code Optimization

Context scaling support for running smaller context models with Claude Code. Scales reported token counts so that auto-compact triggers at the right timing, and SSE keep-alive prevents read timeouts during long prefill.

Multi-Model Serving

Load LLMs, VLMs, embedding models, and rerankers within the same server. Models are managed through a combination of automatic and manual controls:

  • LRU eviction: Least-recently-used models are evicted automatically when memory runs low.
  • Manual load/unload: Interactive status badges in the admin panel let you load or unload models on demand.
  • Model pinning: Pin frequently used models to keep them always loaded.
  • Per-model TTL: Set an idle timeout per model to auto-unload after a period of inactivity.
  • Process memory enforcement: Total memory limit (default: system RAM - 8GB) prevents system-wide OOM.

Per-Model Settings

Configure sampling parameters, chat template kwargs, TTL, model alias, model type override, and more per model directly from the admin panel. Changes apply immediately without server restart.

  • Model alias: set a custom API-visible name. /v1/models returns the alias, and requests accept both the alias and directory name.
  • Model type override: manually set a model as LLM or VLM regardless of auto-detection.
  • Profiles: save named bundles of per-model settings and switch between them from the admin panel. A profile can optionally be exposed as its own model: /v1/models then also lists <model>:<profile> (e.g. qwen3-8b:thinking), which serves on the same engine as the base model with the profile's settings overlaid per request — no extra memory, no reload. When the base model has an alias, the exposed ID is advertised as <alias>:<profile>; the directory-name form keeps working, just like for the base model.

oMLX Chat Template Kwargs

Built-in Chat

Chat directly with any loaded model from the admin panel. Supports conversation history, model switching, dark mode, reasoning model output, and image upload for VLM/OCR models.

oMLX Chat

Model Downloader

Search and download MLX models from HuggingFace directly in the admin dashboard. Browse model cards, check file sizes, and download with one click.

oMLX Model Downloader

Integrations

Set up OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, and Pi directly from the admin dashboard with a single click. No manual config editing required.

oMLX Integrations

Performance Benchmark

One-click benchmarking from the admin panel. Measures prefill (PP) and text generation (TG) tokens per second, with partial prefix cache hit testing for realistic performance numbers.

oMLX Benchmark Tool

macOS Menubar App

Native Swift / SwiftUI menubar app (not Electron). Start, stop, and monitor the server without opening a terminal. Includes persistent serving stats (survives restarts), auto-restart on crash, and Sparkle-driven auto-update.

oMLX Menubar Stats

API Compatibility

Drop-in replacement for OpenAI and Anthropic APIs. Supports streaming usage stats (stream_options.include_usage), Anthropic adaptive thinking, and vision inputs (base64, URL).

Endpoint Description
POST /v1/chat/completions Chat completions (streaming)
POST /v1/completions Text completions (streaming)
POST /v1/messages Anthropic Messages API
POST /v1/embeddings Text embeddings
POST /v1/rerank Document reranking
GET /v1/models List available models

Tool Calling & Structured Output

Supports all function calling formats available in mlx-lm, JSON schema validation, and MCP tool integration. Tool calling requires the model's chat template to support the tools parameter. The following model families are auto-detected via mlx-lm's built-in tool parsers:

Model Family Format
Llama, Qwen, DeepSeek, etc. JSON <tool_call>
Qwen3.5 Series XML <function=...>
Gemma <start_function_call>
GLM (4.7, 5) <arg_key>/<arg_value> XML
MiniMax Namespaced <minimax:tool_call>
Mistral [TOOL_CALLS]
Kimi K2 <|tool_calls_section_begin|>
Longcat <longcat_tool_call>

Models not listed above may still work if their chat template accepts tools and their output uses a recognized <tool_call> XML format. For tool-enabled streaming, assistant text is emitted incrementally while known tool-call control markup is suppressed from visible content; structured tool calls are emitted after parsing the completed turn.

Models

Point --model-dir at a directory containing MLX-format model subdirectories. Two-level organization folders (e.g., mlx-community/model-name/) are also supported.

~/models/
├── Step-3.5-Flash-8bit/
├── Qwen3-Coder-Next-8bit/
├── gpt-oss-120b-MXFP4-Q8/
├── Qwen3.5-122B-A10B-4bit/
└── bge-m3/

Models are auto-detected by type. You can also download models directly from the admin dashboard.

Type Models
LLM Any model supported by mlx-lm
VLM Qwen3.5 Series, GLM-4V, Pixtral, and other mlx-vlm models
OCR DeepSeek-OCR, DOTS-OCR, GLM-OCR
Embedding BERT, BGE-M3, ModernBERT
Reranker ModernBERT, XLM-RoBERTa

CLI Configuration

# Managed background server (macOS app or Homebrew install)
omlx start
omlx stop
omlx restart

# Start with default settings (memory guard tier = balanced, manage via admin UI)
omlx serve --model-dir ~/models

# Choose a memory guard tier at startup
omlx serve --model-dir ~/models --memory-guard safe

# Set a custom memory guard ceiling in GB
omlx serve --model-dir ~/models --memory-guard-gb 48

# Enable SSD cache for KV blocks
omlx serve --model-dir ~/models --paged-ssd-cache-dir ~/.omlx/cache

# Set in-memory hot cache size
omlx serve --model-dir ~/models --hot-cache-max-size 20%

# Adjust max concurrent requests (default: 8)
omlx serve --model-dir ~/models --max-concurrent-requests 16

# With MCP tools
omlx serve --model-dir ~/models --mcp-config mcp.json

# HuggingFace mirror endpoint (for restricted regions)
omlx serve --model-dir ~/models --hf-endpoint https://hf-mirror.com

# API key authentication
omlx serve --model-dir ~/models --api-key your-secret-key
# Localhost-only: skip verification via admin panel global settings

All settings can also be configured from the web admin panel at /admin. Settings are persisted to ~/.omlx/settings.json, and CLI flags take precedence.

Architecture
FastAPI Server (OpenAI / Anthropic API)
    │
    ├── EnginePool (multi-model, LRU eviction, TTL, manual load/unload)
    │   ├── BatchedEngine (LLMs, continuous batching)
    │   ├── VLMEngine (vision-language models)
    │   ├── EmbeddingEngine
    │   └── RerankerEngine
    │
    ├── ProcessMemoryEnforcer (total memory limit, TTL checks)
    │
    ├── Scheduler (FCFS, configurable concurrency)
    │   └── mlx-lm BatchGenerator
    │
    └── Cache Stack
        ├── PagedCacheManager (GPU, block-based, CoW, prefix sharing)
        ├── Hot Cache (in-memory tier, write-back)
        └── PagedSSDCacheManager (SSD cold tier, safetensors format)

Development

CLI Server

git clone https://github.com/jundot/omlx.git
cd omlx
pip install -e ".[dev]"
pytest -m "not slow"

macOS App

The native SwiftUI app lives at apps/omlx-mac/. Requires Xcode 26.5+ and Python 3.11+. venvstacks is declared as a dev dependency so pip install -e ".[dev]" (or uv sync --dev) brings the pinned version in. The build script also falls back to uvx venvstacks or pipx run venvstacks if you prefer a host-global tool runner.

# Stage a runnable oMLX.app (xcodebuild + venvstacks Python layers + ad-hoc sign)
apps/omlx-mac/Scripts/build.sh release

# Result lands at apps/omlx-mac/build/Stage/oMLX.app
open apps/omlx-mac/build/Stage/oMLX.app

# Force a fresh venvstacks rebuild (otherwise it's cached by fingerprint)
apps/omlx-mac/Scripts/build.sh release --rebuild-donor

# Stage with optional GLM-5.2 / MiniMax M3 native custom kernels
apps/omlx-mac/Scripts/build.sh release --with-custom-kernel

First cold build takes 10–20 minutes (venvstacks Python layer assembly). Subsequent builds reuse the cached packaging/_export/ and finish in about 4 minutes. See packaging/README.md for the layer configuration and apps/omlx-mac/ for the Swift sources.

Contributing

Contributions are welcome! See Contributing Guide for details.

  • Bug fixes and improvements
  • Performance optimizations
  • Documentation improvements

License

Apache 2.0

Acknowledgments

  • MLX and mlx-lm by Apple
  • mlx-vlm - Vision-language model inference on Apple Silicon
  • vllm-mlx - oMLX started from vllm-mlx v0.1.0 and evolved significantly with multi-model serving, tiered KV caching, VLM with full paged cache support, an admin panel, and a macOS menu bar app
  • venvstacks - Portable Python environment layering for the macOS app bundle
  • mlx-embeddings - Embedding model support for Apple Silicon
  • dflash-mlx - Block diffusion speculative decoding on Apple Silicon
  • MTPLX - Lightning MTP's verify-shape Metal kernels are powered by MTPLX by Youssof Altoukhi, which also inspired the depth-k pipeline
  • SiliconScope - The menu bar statistics take their design and rendering approach from SiliconScope by Kennt Kim, which also inspired the energy-efficient re-render gating
View on GitHub

Recent activity

commits and pull requests

Releases and announcements

109 total
  1. 0.5.7v0.5.7Aug 4, 20266.4K downloads

    # oMLX 0.5.7 > **Hotfix history since 0.5.4** > > - **0.5.5:** DeepSeek V4 prompt-tail visibility, tiered-cache compatibility, and sparse-prefill fallback cache safety. > - **0.5.6:** DeepSeek V4 CacheList and MTP/PoolingCache integrity, retained-reasoning prompt consistency, long-context indexer safety, and macOS idle-timeout controls. > - **0.5.7:** Official DeepSeek V4 Flash 0731 prompt encoding and safe Claude Code tail-system handling. > > Sorry for the unusually frequent hotfix releases since 0.5.4. DeepSeek V4 Flash 0731 exposed several interacting edge cases across prompt encoding, MTP, PoolingCache, tiered cache, and long-context inference. These updates were necessary to complete and stabilize DeepSeek V4 support. Thank you for your patience and understanding while this work was completed. oMLX 0.5.7 rolls up every fix shipped in 0.5.5 and 0.5.6 and completes the DeepSeek V4 Flash 0731 prompt-format integration. Users upgrading from 0.5.4 can move directly to 0.5.7. ## 0.5.7 Highlights - **DeepSeek V4 now uses the official Flash 0731 reference encoder.** Tool schemas, DSML tool calls, `<tool_result>` history, thinking mode, reasoning effort, role transit

  2. 0.5.6v0.5.6Aug 4, 20261.3K downloads

    > **Hotfix release:** Addresses recurring DeepSeek V4 CacheList signature mismatches and MTP/PoolingCache boundary integrity (#2493, #2500), retained-reasoning prompt consistency (#2501), long-context indexer fallback safety (#2502), and macOS idle-timeout controls (#2498). ## Highlights from 0.5.4 - **DeepSeek V4 Flash 0731 gains DSpark Lightning MTP with up to 85.6% faster code decoding.** Embedded DSpark weights, oQ/oQe, and Metal kernels are supported with matching greedy output. (#2460) - **Inkling Small is supported and accelerated with 1.18–1.23x MTP speedups.** Text and vision serving, oQ/oQe, composite caches, and built-in eight-depth Lightning MTP are included. Community layouts also load; audio is not yet supported. (#2438, #2463; reported by @studioburnside in #2451) - **The context benchmark measures usable context on the current Mac.** It verifies the largest successful prefill and can apply the result through the API, web dashboard, or macOS app. (#2390, #2391) - **Accelerated benchmark runs reach the leaderboard.** MTP, TurboQuant, DFlash, SpecPrefill, and VLM MTP results now upload with feature flags, exact model names, and per-run host metrics. - **MTP benchmark

  3. 0.5.5v0.5.5Aug 3, 20262.7K downloads

    > **Hotfix release:** Addresses DeepSeek V4 prompt-tail visibility (#2490), tiered-cache compatibility (#2487), and sparse-prefill fallback cache safety (#2484). ## Highlights from 0.5.4 - **DeepSeek V4 Flash 0731 gains DSpark Lightning MTP with up to 85.6% faster code decoding.** Embedded DSpark weights, oQ/oQe, and Metal kernels are supported with matching greedy output. (#2460) - **Inkling Small is supported and accelerated with 1.18–1.23x MTP speedups.** Text and vision serving, oQ/oQe, composite caches, and built-in eight-depth Lightning MTP are included. Community layouts also load; audio is not yet supported. (#2438, #2463; reported by @studioburnside in #2451) - **The context benchmark measures usable context on the current Mac.** It verifies the largest successful prefill and can apply the result through the API, web dashboard, or macOS app. (#2390, #2391) - **Accelerated benchmark runs reach the leaderboard.** MTP, TurboQuant, DFlash, SpecPrefill, and VLM MTP results now upload with feature flags, exact model names, and per-run host metrics. - **MTP benchmarks can choose code or novel contexts.** Selectable lengths and bundled English, Japanese, and Korean corpora make

  4. 0.5.4v0.5.4Aug 2, 20263.7K downloads

    oMLX 0.5.4 adds native support and acceleration for DeepSeek V4 Flash 0731, Inkling Small, Step-3.7-Flash, MiMo V2.5, and Laguna S-2.1. It also expands Lightning MTP, DFlash, SpecPrefill, context benchmarking, and long-context cache reliability across the server, web dashboard, and macOS app. ## Highlights - **DeepSeek V4 Flash 0731 gains DSpark Lightning MTP with up to 85.6% faster code decoding.** Embedded DSpark weights, oQ/oQe, and Metal kernels are supported with matching greedy output. (#2460) - **Inkling Small is supported and accelerated with 1.18–1.23x MTP speedups.** Text and vision serving, oQ/oQe, composite caches, and built-in eight-depth Lightning MTP are included. Community layouts also load; audio is not yet supported. (#2438, #2463; reported by @studioburnside in #2451) - **The context benchmark measures usable context on the current Mac.** It verifies the largest successful prefill and can apply the result through the API, web dashboard, or macOS app. (#2390, #2391) - **Accelerated benchmark runs reach the leaderboard.** MTP, TurboQuant, DFlash, SpecPrefill, and VLM MTP results now upload with feature flags, exact model names, and per-run host metrics. - **MTP b

  5. 0.5.4rc2v0.5.4rc2Aug 1, 20261.2K downloads

    This second release candidate adds support for two newly released models, **DeepSeek V4 Flash 0731** with embedded **DSpark Lightning MTP** and Thinking Machines **Inkling Small**, plus **Lightning MTP for Step-3.7-Flash**, kernel-level Inkling decode work, and benchmark uploads for accelerated runs. Everything from [0.5.4rc1](https://github.com/jundot/omlx/releases/tag/v0.5.4rc1) is included, and the notes below only cover what changed since then. ## Highlights since rc1 - **DeepSeek V4 Flash 0731 with DSpark Lightning MTP.** The 0731 checkpoints load natively with their embedded DSpark weights, and the existing Lightning MTP toggle picks the DSpark backend automatically when the checkpoint carries `dspark_*` configuration. oQ/oQe quantize the embedded `mtp.0` through `mtp.2` weights with expert importance collection, and Metal kernels cover the DSpark projections, routed experts, DSA indexer scoring, and block verification. (#2460) M3 Ultra 512 GB, `DeepSeek-V4-Flash-0731-oQ4e-mtp` at 4.36 bpw, single stream, TG128, MTP off, prefix and SSD cache disabled. | Context | 1K | 4K | 8K | 16K | 32K | 64K | 128K | |---|---:|---:|---:|---:|---:|---:|---:| | Prefill tok/s | 531

Code frequency

additions and deletions
+133.6K-133.6KWeek of 2026-02-08: +64,218 linesWeek of 2026-02-08: -267 linesWeek of 2026-02-15: +2,845 linesWeek of 2026-02-15: -586 linesWeek of 2026-02-22: +18,391 linesWeek of 2026-02-22: -5,482 linesWeek of 2026-03-01: +13,702 linesWeek of 2026-03-01: -1,997 linesWeek of 2026-03-08: +19,718 linesWeek of 2026-03-08: -2,244 linesWeek of 2026-03-15: +54,939 linesWeek of 2026-03-15: -7,919 linesWeek of 2026-03-22: +69,334 linesWeek of 2026-03-22: -3,287 linesWeek of 2026-03-29: +9,505 linesWeek of 2026-03-29: -6,331 linesWeek of 2026-04-05: +10,368 linesWeek of 2026-04-05: -1,658 linesWeek of 2026-04-12: +7,180 linesWeek of 2026-04-12: -1,565 linesWeek of 2026-04-19: +44,044 linesWeek of 2026-04-19: -1,022 linesWeek of 2026-04-26: +5,347 linesWeek of 2026-04-26: -775 linesWeek of 2026-05-03: +18,732 linesWeek of 2026-05-03: -2,926 linesWeek of 2026-05-10: +18,922 linesWeek of 2026-05-10: -3,529 linesWeek of 2026-05-17: +9,313 linesWeek of 2026-05-17: -1,843 linesWeek of 2026-05-24: +59,411 linesWeek of 2026-05-24: -12,897 linesWeek of 2026-05-31: +30,766 linesWeek of 2026-05-31: -11,194 linesWeek of 2026-06-07: +21,373 linesWeek of 2026-06-07: -3,662 linesWeek of 2026-06-14: +18,173 linesWeek of 2026-06-14: -4,377 linesWeek of 2026-06-21: +24,914 linesWeek of 2026-06-21: -5,442 linesWeek of 2026-06-28: +8,386 linesWeek of 2026-06-28: -774 linesWeek of 2026-07-05: +34,531 linesWeek of 2026-07-05: -5,116 linesWeek of 2026-07-12: +8,455 linesWeek of 2026-07-12: -1,138 linesWeek of 2026-07-19: +30,841 linesWeek of 2026-07-19: -1,759 linesWeek of 2026-07-26: +60,288 linesWeek of 2026-07-26: -3,161 linesWeek of 2026-08-02: +133,637 linesWeek of 2026-08-02: -23,099 linesFeb 8, 2026Aug 2, 2026
+797.3K lines added, -114.1K removed over the last year.

Commits per week

last 52 weeks
1570Week of 2025-08-09: 0 commitsWeek of 2025-08-16: 0 commitsWeek of 2025-08-23: 0 commitsWeek of 2025-08-30: 0 commitsWeek of 2025-09-06: 0 commitsWeek of 2025-09-13: 0 commitsWeek of 2025-09-20: 0 commitsWeek of 2025-09-27: 0 commitsWeek of 2025-10-04: 0 commitsWeek of 2025-10-11: 0 commitsWeek of 2025-10-18: 0 commitsWeek of 2025-10-25: 0 commitsWeek of 2025-11-01: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 0 commitsWeek of 2026-02-08: 19 commitsWeek of 2026-02-15: 21 commitsWeek of 2026-02-22: 72 commitsWeek of 2026-03-01: 110 commitsWeek of 2026-03-08: 137 commitsWeek of 2026-03-15: 137 commitsWeek of 2026-03-22: 102 commitsWeek of 2026-03-29: 98 commitsWeek of 2026-04-05: 85 commitsWeek of 2026-04-12: 63 commitsWeek of 2026-04-19: 48 commitsWeek of 2026-04-26: 27 commitsWeek of 2026-05-03: 69 commitsWeek of 2026-05-10: 97 commitsWeek of 2026-05-17: 70 commitsWeek of 2026-05-24: 107 commitsWeek of 2026-05-31: 157 commitsWeek of 2026-06-07: 121 commitsWeek of 2026-06-14: 45 commitsWeek of 2026-06-21: 29 commitsWeek of 2026-06-28: 36 commitsWeek of 2026-07-05: 85 commitsWeek of 2026-07-12: 53 commitsWeek of 2026-07-19: 86 commitsWeek of 2026-07-26: 87 commitsWeek of 2026-08-02: 42 commitsAug 9, 2025Aug 2, 2026
2K commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 26 commitsSun 1:00 — 22 commitsSun 2:00 — 21 commitsSun 3:00 — 11 commitsSun 4:00 — 8 commitsSun 5:00 — 15 commitsSun 6:00 — 7 commitsSun 7:00 — 2 commitsSun 8:00 — 1 commitsSun 9:00 — 6 commitsSun 10:00 — 6 commitsSun 11:00 — 7 commitsSun 12:00 — 5 commitsSun 13:00 — 2 commitsSun 14:00 — 11 commitsSun 15:00 — 16 commitsSun 16:00 — 15 commitsSun 17:00 — 18 commitsSun 18:00 — 14 commitsSun 19:00 — 16 commitsSun 20:00 — 8 commitsSun 21:00 — 8 commitsSun 22:00 — 13 commitsSun 23:00 — 25 commitsMon 0:00 — 35 commitsMon 1:00 — 15 commitsMon 2:00 — 12 commitsMon 3:00 — 6 commitsMon 4:00 — 5 commitsMon 5:00 — 3 commitsMon 6:00 — 0 commitsMon 7:00 — 1 commitsMon 8:00 — 1 commitsMon 9:00 — 2 commitsMon 10:00 — 13 commitsMon 11:00 — 6 commitsMon 12:00 — 7 commitsMon 13:00 — 12 commitsMon 14:00 — 13 commitsMon 15:00 — 9 commitsMon 16:00 — 20 commitsMon 17:00 — 14 commitsMon 18:00 — 17 commitsMon 19:00 — 9 commitsMon 20:00 — 5 commitsMon 21:00 — 9 commitsMon 22:00 — 23 commitsMon 23:00 — 19 commitsTue 0:00 — 23 commitsTue 1:00 — 23 commitsTue 2:00 — 6 commitsTue 3:00 — 8 commitsTue 4:00 — 11 commitsTue 5:00 — 4 commitsTue 6:00 — 4 commitsTue 7:00 — 3 commitsTue 8:00 — 6 commitsTue 9:00 — 9 commitsTue 10:00 — 20 commitsTue 11:00 — 35 commitsTue 12:00 — 19 commitsTue 13:00 — 20 commitsTue 14:00 — 19 commitsTue 15:00 — 35 commitsTue 16:00 — 25 commitsTue 17:00 — 31 commitsTue 18:00 — 25 commitsTue 19:00 — 13 commitsTue 20:00 — 7 commitsTue 21:00 — 10 commitsTue 22:00 — 19 commitsTue 23:00 — 24 commitsWed 0:00 — 46 commitsWed 1:00 — 33 commitsWed 2:00 — 3 commitsWed 3:00 — 2 commitsWed 4:00 — 3 commitsWed 5:00 — 5 commitsWed 6:00 — 5 commitsWed 7:00 — 5 commitsWed 8:00 — 8 commitsWed 9:00 — 13 commitsWed 10:00 — 20 commitsWed 11:00 — 14 commitsWed 12:00 — 8 commitsWed 13:00 — 13 commitsWed 14:00 — 13 commitsWed 15:00 — 11 commitsWed 16:00 — 17 commitsWed 17:00 — 14 commitsWed 18:00 — 21 commitsWed 19:00 — 8 commitsWed 20:00 — 3 commitsWed 21:00 — 18 commitsWed 22:00 — 18 commitsWed 23:00 — 34 commitsThu 0:00 — 25 commitsThu 1:00 — 15 commitsThu 2:00 — 9 commitsThu 3:00 — 3 commitsThu 4:00 — 3 commitsThu 5:00 — 1 commitsThu 6:00 — 0 commitsThu 7:00 — 1 commitsThu 8:00 — 12 commitsThu 9:00 — 10 commitsThu 10:00 — 15 commitsThu 11:00 — 20 commitsThu 12:00 — 5 commitsThu 13:00 — 15 commitsThu 14:00 — 13 commitsThu 15:00 — 17 commitsThu 16:00 — 14 commitsThu 17:00 — 13 commitsThu 18:00 — 24 commitsThu 19:00 — 13 commitsThu 20:00 — 3 commitsThu 21:00 — 4 commitsThu 22:00 — 5 commitsThu 23:00 — 17 commitsFri 0:00 — 26 commitsFri 1:00 — 12 commitsFri 2:00 — 12 commitsFri 3:00 — 16 commitsFri 4:00 — 2 commitsFri 5:00 — 9 commitsFri 6:00 — 4 commitsFri 7:00 — 5 commitsFri 8:00 — 6 commitsFri 9:00 — 16 commitsFri 10:00 — 20 commitsFri 11:00 — 20 commitsFri 12:00 — 11 commitsFri 13:00 — 5 commitsFri 14:00 — 8 commitsFri 15:00 — 15 commitsFri 16:00 — 20 commitsFri 17:00 — 3 commitsFri 18:00 — 10 commitsFri 19:00 — 8 commitsFri 20:00 — 11 commitsFri 21:00 — 4 commitsFri 22:00 — 5 commitsFri 23:00 — 11 commitsSat 0:00 — 13 commitsSat 1:00 — 18 commitsSat 2:00 — 13 commitsSat 3:00 — 8 commitsSat 4:00 — 11 commitsSat 5:00 — 3 commitsSat 6:00 — 6 commitsSat 7:00 — 9 commitsSat 8:00 — 10 commitsSat 9:00 — 8 commitsSat 10:00 — 12 commitsSat 11:00 — 6 commitsSat 12:00 — 10 commitsSat 13:00 — 2 commitsSat 14:00 — 11 commitsSat 15:00 — 11 commitsSat 16:00 — 9 commitsSat 17:00 — 13 commitsSat 18:00 — 11 commitsSat 19:00 — 5 commitsSat 20:00 — 2 commitsSat 21:00 — 6 commitsSat 22:00 — 6 commitsSat 23:00 — 11 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.

Who is committing

last 52 weeks
Maintainer commits128 (6%)
Community commits2,015 (94%)

2,143 commits in total over the last year.

DateListRankStars gained
Mar 11, 2026daily#22+113
Mar 10, 2026daily#8+264
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python

  • practical-tutorials/project-based-learning

    Curated list of project-based tutorials

    277.2K stars · Python

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    238.5K stars · JavaScript

  • affaan-m/ECC

    The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

    234.7K stars · JavaScript

  • NousResearch/hermes-agent

    The agent that grows with you

    227K stars · Python