huggingface/speech-to-speechPublic

Build local voice agents with open-source models

AI summary: A modular, open-source voice pipeline for building real-time, low-latency conversational AI agents.

Stars
11.5K
+173 today
Forks
1.4K
Watchers
108
Open issues
87
Open PRs
45
Contributors
~37
Commits
733
Branches
47

PythonApache-2.0Created Aug 7, 2024Last push todayLatest release v0.2.12+1.7K stars this week+3.7K this month

Star history

since Sep 21, 2025
05K10KSep 2025Jan 2026Apr 2026Aug 2026
11.5K stars as of Aug 7, 2026, tracked back to Sep 21, 2025. Historical curve reconstructed from public GitHub event archives, calibrated to the current total.

Contribution activity

commits per day, last 52 weeks
AugSepOctNovDecJanFebMarAprMayJunJulAugMonWedFri2025-08-10: 0 commits2025-08-11: 0 commits2025-08-12: 0 commits2025-08-13: 0 commits2025-08-14: 0 commits2025-08-15: 0 commits2025-08-16: 0 commits2025-08-17: 0 commits2025-08-18: 0 commits2025-08-19: 0 commits2025-08-20: 0 commits2025-08-21: 0 commits2025-08-22: 0 commits2025-08-23: 0 commits2025-08-24: 0 commits2025-08-25: 0 commits2025-08-26: 0 commits2025-08-27: 0 commits2025-08-28: 0 commits2025-08-29: 0 commits2025-08-30: 0 commits2025-08-31: 0 commits2025-09-01: 0 commits2025-09-02: 0 commits2025-09-03: 0 commits2025-09-04: 0 commits2025-09-05: 0 commits2025-09-06: 0 commits2025-09-07: 0 commits2025-09-08: 0 commits2025-09-09: 0 commits2025-09-10: 0 commits2025-09-11: 0 commits2025-09-12: 0 commits2025-09-13: 0 commits2025-09-14: 0 commits2025-09-15: 0 commits2025-09-16: 0 commits2025-09-17: 0 commits2025-09-18: 0 commits2025-09-19: 0 commits2025-09-20: 0 commits2025-09-21: 0 commits2025-09-22: 0 commits2025-09-23: 0 commits2025-09-24: 0 commits2025-09-25: 0 commits2025-09-26: 0 commits2025-09-27: 0 commits2025-09-28: 0 commits2025-09-29: 0 commits2025-09-30: 0 commits2025-10-01: 0 commits2025-10-02: 0 commits2025-10-03: 0 commits2025-10-04: 0 commits2025-10-05: 0 commits2025-10-06: 0 commits2025-10-07: 0 commits2025-10-08: 0 commits2025-10-09: 0 commits2025-10-10: 0 commits2025-10-11: 0 commits2025-10-12: 0 commits2025-10-13: 0 commits2025-10-14: 0 commits2025-10-15: 0 commits2025-10-16: 0 commits2025-10-17: 0 commits2025-10-18: 0 commits2025-10-19: 0 commits2025-10-20: 0 commits2025-10-21: 0 commits2025-10-22: 0 commits2025-10-23: 0 commits2025-10-24: 0 commits2025-10-25: 0 commits2025-10-26: 0 commits2025-10-27: 0 commits2025-10-28: 0 commits2025-10-29: 0 commits2025-10-30: 0 commits2025-10-31: 0 commits2025-11-01: 0 commits2025-11-02: 0 commits2025-11-03: 0 commits2025-11-04: 0 commits2025-11-05: 0 commits2025-11-06: 0 commits2025-11-07: 0 commits2025-11-08: 0 commits2025-11-09: 0 commits2025-11-10: 0 commits2025-11-11: 0 commits2025-11-12: 0 commits2025-11-13: 0 commits2025-11-14: 0 commits2025-11-15: 0 commits2025-11-16: 0 commits2025-11-17: 0 commits2025-11-18: 0 commits2025-11-19: 0 commits2025-11-20: 0 commits2025-11-21: 0 commits2025-11-22: 0 commits2025-11-23: 0 commits2025-11-24: 0 commits2025-11-25: 0 commits2025-11-26: 0 commits2025-11-27: 0 commits2025-11-28: 0 commits2025-11-29: 0 commits2025-11-30: 0 commits2025-12-01: 0 commits2025-12-02: 0 commits2025-12-03: 0 commits2025-12-04: 0 commits2025-12-05: 0 commits2025-12-06: 0 commits2025-12-07: 0 commits2025-12-08: 0 commits2025-12-09: 0 commits2025-12-10: 0 commits2025-12-11: 0 commits2025-12-12: 0 commits2025-12-13: 0 commits2025-12-14: 0 commits2025-12-15: 0 commits2025-12-16: 0 commits2025-12-17: 0 commits2025-12-18: 0 commits2025-12-19: 0 commits2025-12-20: 0 commits2025-12-21: 0 commits2025-12-22: 0 commits2025-12-23: 0 commits2025-12-24: 0 commits2025-12-25: 0 commits2025-12-26: 0 commits2025-12-27: 0 commits2025-12-28: 0 commits2025-12-29: 0 commits2025-12-30: 0 commits2025-12-31: 0 commits2026-01-01: 0 commits2026-01-02: 0 commits2026-01-03: 0 commits2026-01-04: 0 commits2026-01-05: 0 commits2026-01-06: 0 commits2026-01-07: 0 commits2026-01-08: 0 commits2026-01-09: 0 commits2026-01-10: 0 commits2026-01-11: 0 commits2026-01-12: 0 commits2026-01-13: 0 commits2026-01-14: 0 commits2026-01-15: 0 commits2026-01-16: 0 commits2026-01-17: 0 commits2026-01-18: 0 commits2026-01-19: 0 commits2026-01-20: 0 commits2026-01-21: 0 commits2026-01-22: 0 commits2026-01-23: 0 commits2026-01-24: 0 commits2026-01-25: 0 commits2026-01-26: 0 commits2026-01-27: 0 commits2026-01-28: 0 commits2026-01-29: 0 commits2026-01-30: 0 commits2026-01-31: 0 commits2026-02-01: 0 commits2026-02-02: 0 commits2026-02-03: 0 commits2026-02-04: 0 commits2026-02-05: 1 commit2026-02-06: 5 commits2026-02-07: 0 commits2026-02-08: 0 commits2026-02-09: 2 commits2026-02-10: 2 commits2026-02-11: 0 commits2026-02-12: 2 commits2026-02-13: 4 commits2026-02-14: 0 commits2026-02-15: 0 commits2026-02-16: 2 commits2026-02-17: 2 commits2026-02-18: 9 commits2026-02-19: 1 commit2026-02-20: 7 commits2026-02-21: 0 commits2026-02-22: 0 commits2026-02-23: 23 commits2026-02-24: 12 commits2026-02-25: 0 commits2026-02-26: 0 commits2026-02-27: 0 commits2026-02-28: 0 commits2026-03-01: 0 commits2026-03-02: 0 commits2026-03-03: 5 commits2026-03-04: 0 commits2026-03-05: 12 commits2026-03-06: 5 commits2026-03-07: 0 commits2026-03-08: 0 commits2026-03-09: 5 commits2026-03-10: 0 commits2026-03-11: 8 commits2026-03-12: 11 commits2026-03-13: 1 commit2026-03-14: 0 commits2026-03-15: 0 commits2026-03-16: 0 commits2026-03-17: 0 commits2026-03-18: 0 commits2026-03-19: 0 commits2026-03-20: 4 commits2026-03-21: 0 commits2026-03-22: 0 commits2026-03-23: 2 commits2026-03-24: 2 commits2026-03-25: 0 commits2026-03-26: 0 commits2026-03-27: 0 commits2026-03-28: 0 commits2026-03-29: 0 commits2026-03-30: 0 commits2026-03-31: 0 commits2026-04-01: 0 commits2026-04-02: 0 commits2026-04-03: 0 commits2026-04-04: 0 commits2026-04-05: 0 commits2026-04-06: 0 commits2026-04-07: 3 commits2026-04-08: 1 commit2026-04-09: 0 commits2026-04-10: 1 commit2026-04-11: 0 commits2026-04-12: 0 commits2026-04-13: 0 commits2026-04-14: 1 commit2026-04-15: 5 commits2026-04-16: 11 commits2026-04-17: 3 commits2026-04-18: 0 commits2026-04-19: 0 commits2026-04-20: 3 commits2026-04-21: 1 commit2026-04-22: 2 commits2026-04-23: 2 commits2026-04-24: 1 commit2026-04-25: 0 commits2026-04-26: 0 commits2026-04-27: 1 commit2026-04-28: 1 commit2026-04-29: 3 commits2026-04-30: 9 commits2026-05-01: 2 commits2026-05-02: 0 commits2026-05-03: 0 commits2026-05-04: 21 commits2026-05-05: 9 commits2026-05-06: 0 commits2026-05-07: 3 commits2026-05-08: 19 commits2026-05-09: 0 commits2026-05-10: 0 commits2026-05-11: 0 commits2026-05-12: 0 commits2026-05-13: 0 commits2026-05-14: 0 commits2026-05-15: 0 commits2026-05-16: 0 commits2026-05-17: 0 commits2026-05-18: 0 commits2026-05-19: 0 commits2026-05-20: 1 commit2026-05-21: 0 commits2026-05-22: 0 commits2026-05-23: 0 commits2026-05-24: 0 commits2026-05-25: 0 commits2026-05-26: 5 commits2026-05-27: 11 commits2026-05-28: 4 commits2026-05-29: 1 commit2026-05-30: 1 commit2026-05-31: 0 commits2026-06-01: 0 commits2026-06-02: 0 commits2026-06-03: 0 commits2026-06-04: 1 commit2026-06-05: 3 commits2026-06-06: 1 commit2026-06-07: 0 commits2026-06-08: 1 commit2026-06-09: 14 commits2026-06-10: 1 commit2026-06-11: 3 commits2026-06-12: 1 commit2026-06-13: 0 commits2026-06-14: 0 commits2026-06-15: 0 commits2026-06-16: 0 commits2026-06-17: 0 commits2026-06-18: 1 commit2026-06-19: 0 commits2026-06-20: 0 commits2026-06-21: 0 commits2026-06-22: 1 commit2026-06-23: 1 commit2026-06-24: 0 commits2026-06-25: 0 commits2026-06-26: 1 commit2026-06-27: 0 commits2026-06-28: 0 commits2026-06-29: 2 commits2026-06-30: 1 commit2026-07-01: 1 commit2026-07-02: 0 commits2026-07-03: 10 commits2026-07-04: 0 commits2026-07-05: 0 commits2026-07-06: 0 commits2026-07-07: 1 commit2026-07-08: 0 commits2026-07-09: 1 commit2026-07-10: 0 commits2026-07-11: 0 commits2026-07-12: 1 commit2026-07-13: 5 commits2026-07-14: 0 commits2026-07-15: 6 commits2026-07-16: 0 commits2026-07-17: 3 commits2026-07-18: 0 commits2026-07-19: 0 commits2026-07-20: 1 commit2026-07-21: 0 commits2026-07-22: 0 commits2026-07-23: 0 commits2026-07-24: 0 commits2026-07-25: 0 commits2026-07-26: 0 commits2026-07-27: 0 commits2026-07-28: 2 commits2026-07-29: 5 commits2026-07-30: 10 commits2026-07-31: 10 commits2026-08-01: 1 commit2026-08-02: 1 commit2026-08-03: 5 commits2026-08-04: 13 commits2026-08-05: 10 commits2026-08-06: 11 commits2026-08-07: 0 commits2026-08-08: 0 commits
381 commits in the last yearLessMore

Signals and awards

derived from tracked data
  • Widely adopted

    11,547 stars

  • Rising fast

    +1,728 stars this week

  • Actively maintained

    Pushed within 48 hours

  • Permissive license

    Apache-2.0

  • Continuous integration

    Automated checks passing

  • Repeat trending

    13 trending appearances

What speech-to-speech does

This project provides a complete, thread-based cascade for real-time speech interactions: Voice Activity Detection (VAD) using Silero, Speech-to-Text (STT) via models like Whisper or Parakeet, an LLM reasoning engine, and Text-to-Speech (TTS) synthesis via Qwen3 or Kokoro. It operates in multiple modes including a local terminal client, a raw TCP socket server, and a WebSocket implementation of the OpenAI Realtime API protocol. The modular architecture allows developers to dynamically hot-swap different Hugging Face models for each pipeline stage via CLI flags, optimizing for either maximum reasoning quality or the lowest possible latency on specific hardware like NVIDIA GPUs or Apple Silicon.

This framework is for AI researchers and application developers who want to build, test, and self-host low-latency voice assistants using open weights and models from the Hugging Face Hub without relying on proprietary cloud voice APIs.

  • OpenAI Realtime Protocol: Exposes a WebSocket server compatible with standard OpenAI Realtime clients for drop-in voice streaming.
  • Hardware-optimized execution: Automatically leverages MPS on Apple Silicon or GGML/CUDA on Linux to minimize latency across the transcription-generation-synthesis cascade.
  • Interchangeable backends: Allows developers to swap TTS (Qwen3, Kokoro, ChatTTS) or STT (Whisper, Parakeet) models seamlessly via CLI arguments.
  • Live transcription: Streams partial speech-to-text transcripts to the client in real-time before the LLM begins generating a response.
  • Local and containerized modes: Provides a direct local microphone/speaker interface as well as a pre-configured Docker Compose stack for remote server deployments.

Where teams use it

Voice AI prototyping

Researchers use the local microphone mode to quickly test the conversational latency of new Hugging Face models on their MacBooks.

Self-hosted voice assistants

Developers deploy the Dockerized pipeline on a home server to power a completely private, offline smart speaker.

API drop-in replacement

Startups point their existing OpenAI Realtime clients to this local server to save API costs while maintaining the same streaming voice features.

Custom hardware optimization

Hardware engineers test specific quantization variants (e.g., 4-bit vs 8-bit MLX models) to find the perfect latency-quality tradeoff for edge devices.

Getting started: pip install speech-to-speech speech-to-speech --local_mac_optimal_settings

README

main branch
 

Speech To Speech: Build voice agents with open-source models

PyPI Python License GitHub Trending: #1 Repository of the Day

A low-latency, fully modular voice-agent pipeline: VAD -> STT -> LLM -> TTS, exposed through an OpenAI Realtime-compatible WebSocket API. Every component is swappable. The LLM slot speaks OpenAI-compatible protocols, so you can point it at a hosted provider, at HF Inference Providers, or at a vLLM or llama.cpp server on your own hardware for a fully local, fully open stack.

This pipeline runs in production as the conversation backend for thousands of Reachy Mini robots.

Switching an OpenAI Realtime client endpoint from hosted OpenAI to a self-hosted speech-to-speech server

Quickstart

pip install speech-to-speech
export OPENAI_API_KEY=...
speech-to-speech serve

This starts an OpenAI Realtime-compatible server at ws://localhost:8765/v1/realtime using Parakeet TDT for local STT, an OpenAI-compatible LLM, and Qwen3-TTS for local speech output.

Talk to it from a second terminal:

speech-to-speech talk --url ws://127.0.0.1:8765/v1/realtime

To start the server and packaged microphone/speaker client in one command:

speech-to-speech local

Prefer to keep the LLM on your own machine? Serve Gemma 4 with llama.cpp:

llama-server -hf ggml-org/gemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-full

Then point the OpenAI-compatible LLM backend at it:

speech-to-speech serve \
    --model_name "ggml-org/gemma-4-E4B-it-GGUF" \
    --responses_api_base_url "http://127.0.0.1:8080/v1" \
    --responses_api_api_key ""

Any OpenAI Realtime-compatible client can connect. See Realtime API for the protocol and LLM backends for provider and local-server options.

Index

How it works

The pipeline is a cascade of four components, each running in its own thread and connected by queues:

  1. Voice Activity Detection (VAD): Silero VAD v5 detects speech boundaries and turn-taking.
  2. Speech to Text (STT): transcribes the user's turn, with optional live partial transcripts.
  3. Language Model (LLM): generates the response, streaming text and tool calls.
  4. Text to Speech (TTS): synthesizes audio and streams it back to the client.

Every stage has multiple interchangeable backends, selected via CLI flags. The code is designed for easy modification, with a focus on models available through Transformers and the Hugging Face Hub.

Installation

Requires Python 3.10+.

pip install speech-to-speech

The default install covers the standard realtime path:

  • Parakeet TDT for STT
  • OpenAI-compatible API for the language model
  • Qwen3-TTS for speech output, using the GGML backend by default on non-macOS platforms and mlx-audio on Apple Silicon
  • local audio and realtime server modes

macOS and non-macOS dependencies are resolved automatically via platform markers in pyproject.toml.

CUDA Note for Qwen3-TTS

On Linux, the Qwen3-TTS GGML backend comes from faster-qwen3-tts[ggml]. Its default qwentts-cpp-python wheel on PyPI targets CUDA 12.8. If your machine does not have the CUDA 12 runtime that wheel expects, install the matching wheel from the Hugging Face wheelhouse before installing speech-to-speech:

# CUDA 13.x
pip install "qwentts-cpp-python==0.3.1+cu130" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu130

# CUDA 12.4
pip install "qwentts-cpp-python==0.3.1+cu124" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cu124

# CPU-only fallback
pip install "qwentts-cpp-python==0.3.1+cpu" \
  -f https://huggingface.co/datasets/andito/qwentts-cpp-python-wheels/tree/main/whl/cpu

pip install speech-to-speech

To use the previous CUDA-graphs implementation instead of GGML, pass --qwen3_tts_backend torch.

Optional Components

Optional components are installed with pip extras:

pip install "speech-to-speech[kokoro]"          # Kokoro-82M TTS on non-macOS
pip install "speech-to-speech[pocket]"          # Pocket TTS
pip install "speech-to-speech[chattts]"         # ChatTTS
pip install "speech-to-speech[faster-whisper]"  # Faster Whisper STT
pip install "speech-to-speech[whisper-mlx]"     # Lightning Whisper MLX STT on macOS
pip install "speech-to-speech[paraformer]"      # Paraformer STT through FunASR
pip install "speech-to-speech[mlx-lm]"          # mlx-vlm support for vision models on macOS

Deprecated implementations, including MeloTTS, live in archive/ and are no longer wired into the CLI.

Note on DeepFilterNet: DeepFilterNet, used for optional audio enhancement in VAD, requires numpy<2 and conflicts with Pocket TTS, which requires numpy>=2. Install it manually only in environments where you are not using Pocket TTS.

From Source

git clone https://github.com/huggingface/speech-to-speech.git
cd speech-to-speech
uv sync

This installs the package in editable mode and makes the speech-to-speech CLI available.

Supported Components

Component Backend Platforms Install
VAD Silero VAD v5 all built-in
STT Parakeet TDT (default) CUDA / CPU through nano-parakeet, Apple Silicon through MLX built-in
STT Whisper through Transformers CUDA / CPU built-in
STT Faster Whisper CUDA / CPU faster-whisper
STT Lightning Whisper MLX Apple Silicon whisper-mlx
STT MLX Audio Whisper Apple Silicon built-in on macOS
STT Paraformer CUDA / CPU paraformer
LLM OpenAI-compatible API (responses-api, chat-completions) hosted providers or self-hosted servers built-in
LLM Transformers CUDA / CPU built-in
LLM mlx-lm Apple Silicon built-in on macOS
TTS Qwen3-TTS (default) GGML / CUDA on Linux, mlx-audio on macOS built-in
TTS Kokoro-82M CUDA / CPU, Apple Silicon kokoro on non-macOS; built-in on macOS
TTS Pocket TTS CPU / CUDA pocket
TTS ChatTTS CUDA / CPU chattts
TTS MMS TTS CUDA / CPU built-in

Select implementations with --stt, --llm_backend, and --tts. Run speech-to-speech serve -h for exact values and backend-specific flags.

Commands

Command Behavior Use it when
serve Runs the pipeline server over OpenAI Realtime WebSocket and WebRTC. You are building an app or device against the API.
talk --url <full-realtime-url> Runs the packaged microphone/speaker client. You want to talk to an existing Realtime server.
local Composes serve and talk in-process over loopback. You want to run the server and talk to it from one command.

serve binds to 127.0.0.1 by default; pass --host 0.0.0.0 explicitly for network exposure. local always binds to loopback and connects the same packaged client at ws://127.0.0.1:<port>/v1/realtime.

Migrating from --mode

--mode is deprecated and will stop working soon. During this migration window, speech-to-speech --mode realtime runs speech-to-speech serve, and speech-to-speech --mode local runs speech-to-speech local; both print a warning. All other mode values have been removed and exit with guidance to use the new commands.

Realtime Server

export OPENAI_API_KEY=...
speech-to-speech serve

This is equivalent to:

speech-to-speech serve \
    --thresh 0.6 \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_model_name Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice \
    --qwen3_tts_speaker Aiden \
    --qwen3_tts_language auto \
    --qwen3_tts_backend ggml \
    --qwen3_tts_non_streaming_mode True \
    --qwen3_tts_mlx_quantization 6bit \
    --model_name gpt-5.4-mini \
    --chat_size 30 \
    --responses_api_stream \
    --enable_live_transcription

The default model is gpt-5.4-mini through the OpenAI Responses API. Override it with --model_name, and set --responses_api_base_url for another OpenAI-compatible provider or server.

Local Mac

speech-to-speech local --mac-optimal-settings

Optionally with a specific LLM:

speech-to-speech local \
    --mac-optimal-settings \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

This setting:

  • Uses MPS defaults for supported model components.
  • Sets Parakeet TDT for STT.
  • Sets MLX LM as the LLM backend.
  • Sets Qwen3-TTS for TTS, using mlx-audio with the 6bit MLX variant by default.

The preset supplies these as defaults only: explicit --device, component-device flags such as --qwen3_tts_device, and --stt, --llm_backend, --model_name, and --tts all win. Use it with serve instead of local when you want to expose the server without starting the microphone/speaker client.

--tts pocket and --tts kokoro are also valid on macOS.

To compare the MLX quantization variants locally:

python scripts/benchmark_tts.py \
    --handlers qwen3 \
    --iterations 3 \
    --qwen3_mlx_quantizations bf16 4bit 6bit 8bit

Docker

Install the NVIDIA Container Toolkit, then:

docker compose up

The compose file starts a llama.cpp server with Gemma 4 and the Realtime server, exposing ports 8080 and 8765.

Realtime API

Realtime mode supports the OpenAI Realtime protocol over WebSocket and WebRTC, with live transcription and low-latency turn-taking. WebSocket clients connect at /v1/realtime:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8765/v1",
    websocket_base_url="ws://localhost:8765/v1",
    api_key="not-needed",
)

with client.realtime.connect(model="local") as conn:
    conn.send(
        {
            "type": "session.update",
            "session": {
                "type": "realtime",
                "instructions": "You are a helpful assistant.",
                "audio": {
                    "input": {
                        "turn_detection": {
                            "type": "server_vad",
                            "interrupt_response": True,
                        }
                    }
                },
            },
        }
    )

    for event in conn:
        print(event.type)

The server implements the core Realtime event set: input_audio_buffer.append, session.update, conversation.item.create, response.create, and response.cancel inbound; speech start/stop, streaming transcription, audio deltas, tool calls, and response.done outbound. The full event reference, architecture, and design details live in the Realtime Engine README.

LLM Proxy

With --enable_llm_proxy, the realtime server also exposes the remote LLM it is configured with as a plain OpenAI compatible endpoint, so a client can run side tasks (summaries, titles, background agents) with tools and streaming, fully concurrent with the voice conversation and never interrupted by new speech:

  • POST /v1/chat/completions when running --llm_backend chat-completions
  • POST /v1/responses when running --llm_backend responses-api

The server performs no authentication and no throttling of its own. Enable the proxy only on a trusted network, or deploy the server behind a gateway that owns access control. The s2s-endpoint compute replica is such a gateway: it opens these paths only to clients that created their session with an HF token, checks the API key against that token, and applies a rate limit per user. Point the stock OpenAI SDK at whichever host you talk to; this server ignores the API key (a gateway in front decides what it must be):

from openai import OpenAI

llm = OpenAI(base_url="http://localhost:8765/v1", api_key="unused")
completion = llm.chat.completions.create(
    model="anything",  # ignored: the server forces its configured --model_name
    messages=[{"role": "user", "content": "Summarize the conversation so far: ..."}],
)

Requests are stateless (send the full message list each time) and are proxied to the configured upstream with the key held by the server, which never reaches clients. The model field is always overwritten with the server configured --model_name. The proxy is off by default, requires a remote backend (chat-completions or responses-api), and answers 501 with the reason otherwise.

LLM Backends

The LLM is the most compute-intensive and highest-latency component in the pipeline. A single forward pass through a large model can dominate end-to-end response time, so choosing the right backend for your hardware and latency budget matters. The pipeline supports:

  • Local inference: transformers on CUDA / CPU and mlx-lm on Apple Silicon.
  • Self-hosted servers: responses-api and chat-completions can point at a local vLLM or llama.cpp server.
  • Provider APIs: the same backends work with OpenAI, HF Inference Providers, OpenRouter, and other OpenAI-compatible providers.

Two API backends are available, sharing the same --responses_api_* connection flags:

  • --llm_backend responses-api (default) targets /v1/responses.
  • --llm_backend chat-completions targets /v1/chat/completions.

Direct Audio Input (No STT)

Use --stt none --llm_backend chat-completions to send each completed VAD audio segment directly to an audio-input model. Direct audio mode is not supported with --llm_backend responses-api: a model may accept audio through /v1/chat/completions without supporting /v1/responses, including OpenAI's gpt-audio-1.5.

You must explicitly set --model_name to a model that accepts audio: the default gpt-5.4-mini accepts text and image input, but not audio. Check the provider's model documentation and endpoint support before enabling this mode. For OpenAI, see the GPT-5.4 mini model card and audio-input guide.

speech-to-speech serve \
    --stt none \
    --llm_backend chat-completions \
    --model_name "YOUR_AUDIO_CAPABLE_MODEL" \
    --responses_api_base_url "https://provider.example/v1" \
    --responses_api_api_key "$PROVIDER_API_KEY"

OpenAI-compatible servers represent input audio differently. Use --responses_api_audio_content_type input_audio (the default) for embedded WAV base64, or --responses_api_audio_content_type audio_url for a base64 data URL.

The examples below pair Parakeet TDT for local STT and Qwen3-TTS for local TTS with different LLM backends.

Responses API Backend

Works with any provider or server that implements the OpenAI Responses API. Point --responses_api_base_url at the endpoint and set --model_name accordingly:

Provider / server --responses_api_base_url --responses_api_api_key
OpenAI omit, uses OpenAI default $OPENAI_API_KEY
HF Inference Providers https://router.huggingface.co/v1 $HF_TOKEN
OpenRouter https://openrouter.ai/api/v1 $OPENROUTER_API_KEY
vLLM http://localhost:8000/v1 omit or any string
llama.cpp http://127.0.0.1:8080/v1 empty string
# OpenAI
speech-to-speech local \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_mlx_quantization 6bit \
    --model_name "gpt-4o-mini" \
    --responses_api_api_key "$OPENAI_API_KEY" \
    --responses_api_stream \
    --enable_live_transcription
# HF Inference Providers: Qwen3.5-9B via Together
speech-to-speech local \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_mlx_quantization 6bit \
    --model_name "Qwen/Qwen3.5-9B:together" \
    --responses_api_base_url "https://router.huggingface.co/v1" \
    --responses_api_api_key "$HF_TOKEN" \
    --responses_api_stream \
    --enable_live_transcription
# HF Inference Providers: GPT-oss-20B via Groq
speech-to-speech serve \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --qwen3_tts_mlx_quantization 6bit \
    --model_name "openai/gpt-oss-20b:groq" \
    --responses_api_base_url "https://router.huggingface.co/v1" \
    --responses_api_api_key "$HF_TOKEN" \
    --responses_api_stream \
    --enable_live_transcription

Chat Completions Backend

Identical configuration to responses-api, reusing the same --responses_api_* connection flags, but talks to /v1/chat/completions instead of /v1/responses. Prefer it when:

  • the provider ignores chat_template_kwargs.enable_thinking on the Responses path and needs a reasoning_effort knob to suppress reasoning, or
  • the server's Responses streaming tool-call path is unreliable, while its Chat Completions tool-call streaming is solid. This is useful for some vLLM builds; see #312.

Add --responses_api_reasoning_effort none to disable reasoning on providers where the chat-template flag has no effect:

# vLLM serving a Qwen model with tool calling
speech-to-speech serve \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "Qwen/Qwen3-4B-Instruct-2507" \
    --responses_api_base_url "http://localhost:8000/v1" \
    --responses_api_stream
# Gemma 4 31B via the HF router on Cerebras, with reasoning disabled for low voice latency
speech-to-speech serve \
    --stt parakeet-tdt \
    --llm_backend chat-completions \
    --tts qwen3 \
    --model_name "google/gemma-4-31B-it:cerebras" \
    --responses_api_base_url "https://router.huggingface.co/v1" \
    --responses_api_api_key "$HF_TOKEN" \
    --responses_api_reasoning_effort none \
    --responses_api_stream

Fully Local

Run the LLM in a separate llama.cpp process for the lowest-friction fully local setup, as shown in the Reachy Mini local conversation guide:

For a fully local native-audio setup with the browser demo, Realtime turn revisions, and barge-in, see the tested Gemma 4 12B speech-to-speech example for Apple Silicon.

# Terminal 1: llama.cpp serving Gemma 4
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF -np 2 -c 65536 -fa on --swa-full
# Terminal 2: speech-to-speech using that local LLM server
speech-to-speech serve \
    --stt parakeet-tdt \
    --llm_backend responses-api \
    --tts qwen3 \
    --model_name "ggml-org/gemma-4-E4B-it-GGUF" \
    --responses_api_base_url "http://127.0.0.1:8080/v1" \
    --responses_api_api_key "" \
    --responses_api_stream \
    --enable_live_transcription

Use speech-to-speech local when you want to run the same server and talk through the machine hosting it. In-process local backends are available with --llm_backend mlx-lm on Apple Silicon or --llm_backend transformers on CUDA / CPU.

Multi-Language Support

Language coverage depends on the STT and TTS backends you pick, not on the pipeline itself:

Component Backend Languages
STT Parakeet TDT (default) 25 European languages
STT Whisper / Whisper MLX / Faster Whisper Broad multilingual coverage, depending on the selected Whisper checkpoint
STT Paraformer Depends on the selected FunASR checkpoint; the default is Chinese-oriented
TTS Qwen3-TTS (default) Multilingual, with --qwen3_tts_language auto by default
TTS Kokoro Multiple language/voice mappings, depending on backend availability
TTS ChatTTS English and Chinese
TTS MMS TTS Broad multilingual coverage through MMS checkpoints

Make sure the STT, LLM, and TTS you pair all cover your target language(s). Two usage patterns:

  • Single language: set --language to the target language code. The default is en.
  • Language switching: set --language auto. The STT detects the language of each spoken prompt and forwards it to the LLM. Optionally add --enable_lang_prompt to append a "Please reply to my message in ..." instruction. It defaults to False; large LLMs usually infer the language from context, but the explicit instruction can help smaller models.

Automatic language detection:

speech-to-speech serve \
    --stt parakeet-tdt \
    --language auto \
    --llm_backend mlx-lm \
    --model_name "mlx-community/Qwen3-4B-Instruct-2507-bf16"

A single non-English language, Chinese in this example:

speech-to-speech serve \
    --stt whisper-mlx \
    --stt_model_name large-v3 \
    --language zh \
    --llm_backend mlx-lm \
    --model_name mlx-community/Qwen3-4B-Instruct-2507-bf16

Both commands also work with --mac-optimal-settings; explicit --stt flags override the defaults it sets.

Pocket TTS

Pocket TTS from Kyutai Labs provides streaming TTS with voice cloning:

speech-to-speech serve \
    --tts pocket \
    --pocket_tts_voice jean \
    --pocket_tts_device cpu

Available voice presets: alba, marius, javert, jean, fantine, cosette, eponine, azelma. Custom voice files and Hugging Face paths also work.

CLI Reference

References for pipeline CLI arguments live in the arguments classes and in speech-to-speech serve -h. Client arguments are listed by speech-to-speech talk -h.

Module-Level Parameters

See ModuleArguments. It allows setting:

  • a common --device, if every part should run on the same device
  • macOS model/device defaults (--mac-optimal-settings)
  • STT implementation (--stt)
  • LLM backend (--llm_backend: transformers, mlx-lm, responses-api, or chat-completions)
  • TTS implementation (--tts)
  • logging level
  • realtime pipeline pool size (--num_pipelines)

VAD Parameters

See VADHandlerArguments. Notable options:

  • --thresh: threshold value to trigger voice activity detection.
  • --min_speech_ms: minimum duration of detected voice activity to be considered speech.
  • --min_speech_continuation_ms: sustain-bar hysteresis threshold for speech that continues a reopenable soft-ended, uncommitted turn within the reopen window. The default and recommended pairing is --min_speech_ms 384 --min_speech_continuation_ms 192.
  • --min_silence_ms: minimum length of silence intervals for segmenting speech. Default is 64 ms.
  • --short_segment_merge_ms: optional merge window for stitching adjacent VAD segments that are each shorter than --min_speech_ms.
  • --speculative_reopen_ms: delay response commitment for 800 ms after a soft-ended turn so immediately resumed speech can reopen it.
  • --unanswered_reopen_ms: sanity cap on how long a soft-ended speculative turn that has not yet received any assistant output stays reopenable. With Smart Turn enabled, this is clamped to at least --smart_turn_max_wait_ms so a turn remains reopenable for its full grace.

Smart Turn endpointing

Smart Turn v3.2 can validate Silero's end-of-speech decisions using the content and prosody of the current turn. Silero finalizes the segment and STT/LLM work may begin speculatively. Complete turns start processing immediately and use --speculative_reopen_ms (800 ms by default) before committing output. Incomplete turns wait --smart_turn_incomplete_delay_ms (600 ms by default) before starting STT/LLM work, while their output remains gated by --smart_turn_max_wait_ms (2 seconds by default). If speech resumes during either delay, the existing turn is reopened as a newer revision, the accumulated audio is re-emitted, and work from the previous revision is discarded before it reaches the user.

The base package includes the quantized CPU runtime and enables Smart Turn by default:

pip install speech-to-speech
speech-to-speech serve

The latest supported v3.2 CPU checkpoint downloads from the Hugging Face Hub on first use. Pass --smart_turn_model_path /path/to/model.onnx to use a local model, or --no_smart_turn to disable Smart Turn. Smart Turn is enabled by default for server sessions and the packaged local client.

Tune the completion cutoff with --smart_turn_threshold (default 0.5). A higher threshold makes ambiguous pauses more likely to use the longer speculative response grace.

STT, LLM, and TTS Parameters

model_name, torch_dtype, and device are exposed for each STT, LLM, and TTS implementation. STT and TTS parameters use the handler prefix, for example --stt_model_name or --qwen3_tts_device. LLM model selection and chat settings are shared across backends via unprefixed flags, for example --model_name and --chat_size; backend-specific flags use the responses_api_ prefix for the responses-api and chat-completions backends and the llm_ prefix for local backends.

For example:

# Local transformers/mlx-lm backend
--model_name google/gemma-2b-it

# OpenAI-compatible backend
--llm_backend responses-api --model_name deepseek-chat --responses_api_base_url https://api.deepseek.com

Generation Parameters

Other generation parameters can be set using the handler prefix plus _gen_, for example --stt_gen_max_new_tokens 128 or --llm_gen_temperature 0.7. Parameters not yet exposed can be added to the relevant arguments class.

Contributing

Issues and PRs are welcome. Good starting points are the open issues. For larger changes, open an issue first to discuss the approach.

For local development:

uv sync
pytest
ruff check

Star History

Star History Chart

Citations

If you use this pipeline, please also cite the component models you run. The defaults are:

Silero VAD

@misc{SileroVAD,
  author = {Silero Team},
  title = {Silero VAD: pre-trained enterprise-grade Voice Activity Detector (VAD), Number Detector and Language Classifier},
  year = {2021},
  publisher = {GitHub},
  journal = {GitHub repository},
  howpublished = {\url{https://github.com/snakers4/silero-vad}},
  email = {hello@silero.ai}
}

Parakeet TDT

@misc{parakeet-tdt,
  author = {NVIDIA},
  title = {Parakeet TDT 0.6B v3},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3}}
}

Qwen3-TTS

@misc{qwen3-tts,
  author = {Qwen Team},
  title = {Qwen3-TTS},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice}}
}

Citations for optional backends such as Kokoro, Pocket TTS, ChatTTS, Whisper variants, Paraformer, and MMS live in the respective component READMEs.

View on GitHub

Recent activity

commits and pull requests

Releases and announcements

4 total
  1. v0.2.12v0.2.12Aug 5, 2026

    `speech-to-speech` 0.2.12 is the final planned release in the 0.2.x line before the next round of larger changes. It brings smarter turn-taking, WebRTC support, direct audio input for audio-capable LLMs, more complete OpenAI Realtime protocol behavior, and a significantly improved browser demo. ### Highlights - **Smarter endpointing with Smart Turn v3.2.** Realtime mode now enables the quantized CPU Smart Turn model by default to distinguish completed turns from mid-thought pauses while speculative STT and LLM work continues. Use `--no_smart_turn` to retain Silero-only endpointing. ([#192](https://github.com/huggingface/speech-to-speech/pull/192)) - **WebRTC transport for the OpenAI Realtime API.** Install the new `webrtc` extra to use SDP negotiation at `POST /v1/realtime/calls`, RTP audio, and the `oai-events` data channel alongside the existing WebSocket transport. Both transports share the same pipeline pool, event dispatch, cancellation, and interruption behavior. ([#352](https://github.com/huggingface/speech-to-speech/pull/352)) - **Direct audio input for audio-capable LLMs.** Run with `--stt none`, the Chat Completions backend, and an explicitly selected a

  2. v0.2.11v0.2.11Aug 3, 2026

    ## What's Changed * Bump the actions group with 2 updates by @dependabot[bot] in https://github.com/huggingface/speech-to-speech/pull/316 * Support text-only and out-of-band (conversation=none) responses by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/318 * Emit assistant transcript for fresh response when discard guard is stuck by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/321 * Add chat-completions LLM backend (OpenAI /v1/chat/completions) by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/322 * Default Qwen3 TTS to GGML by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/325 * Fix two mid-generation conversation races (tool call & image) by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/326 * Defer client conversation items during an active response by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/327 * Update HF router Gemma chat completions example by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/328 * Refresh README docs by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/330 * chore: remove deprecated melo option

  3. v0.2.10v0.2.10Jun 11, 2026

    ## What's Changed * Add --num_pipelines pool for concurrent realtime sessions by @A-Mahla in https://github.com/huggingface/speech-to-speech/pull/282 * Add speculative turn revisions by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/255 * Bump the actions group with 5 updates by @dependabot[bot] in https://github.com/huggingface/speech-to-speech/pull/293 * Normalize Qwen3-TTS language aliases by @kamjin3086 in https://github.com/huggingface/speech-to-speech/pull/300 * [codex] Fix TTS benchmark input message by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/303 * Bump the actions group with 2 updates by @dependabot[bot] in https://github.com/huggingface/speech-to-speech/pull/301 * Fix Paraformer progressive transcription events by @kamjin3086 in https://github.com/huggingface/speech-to-speech/pull/299 * Improve speculative VAD reopen and continuation handling by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/307 * Fix tool-only realtime response completion by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/306 * Shorten voice prompt and improve tool lead-ins by @andimarafioti in https://gi

  4. 2025 release2025Feb 6, 2026

    ## What's Changed * Minor doc fix. by @Vaibhavs10 in https://github.com/huggingface/speech-to-speech/pull/2 * Fix missing sounddevice module by @AlexHayton in https://github.com/huggingface/speech-to-speech/pull/7 * Update README.md by @RodriMora in https://github.com/huggingface/speech-to-speech/pull/23 * fix issue with ntlk by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/29 * Dockerize by @codearranger in https://github.com/huggingface/speech-to-speech/pull/22 * Add support to MPS by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/20 * adding apache license by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/31 * refactor arguments folder + run ruff by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/32 * Allow LM selection and MLX Gemma by @RonanKMcGovern in https://github.com/huggingface/speech-to-speech/pull/40 * Improvements mlx pipeline by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/41 * refactor all the handlers - folder structure by @andimarafioti in https://github.com/huggingface/speech-to-speech/pull/43 * add min new tokens by @andimarafioti

Code frequency

additions and deletions
+14.4K-14.4KWeek of 2025-08-10: +0 linesWeek of 2025-08-10: -0 linesWeek of 2025-08-17: +0 linesWeek of 2025-08-17: -0 linesWeek of 2025-08-24: +0 linesWeek of 2025-08-24: -0 linesWeek of 2025-08-31: +0 linesWeek of 2025-08-31: -0 linesWeek of 2025-09-07: +0 linesWeek of 2025-09-07: -0 linesWeek of 2025-09-14: +0 linesWeek of 2025-09-14: -0 linesWeek of 2025-09-21: +0 linesWeek of 2025-09-21: -0 linesWeek of 2025-09-28: +0 linesWeek of 2025-09-28: -0 linesWeek of 2025-10-05: +0 linesWeek of 2025-10-05: -0 linesWeek of 2025-10-12: +0 linesWeek of 2025-10-12: -0 linesWeek of 2025-10-19: +0 linesWeek of 2025-10-19: -0 linesWeek of 2025-10-26: +0 linesWeek of 2025-10-26: -0 linesWeek of 2025-11-02: +0 linesWeek of 2025-11-02: -0 linesWeek of 2025-11-09: +0 linesWeek of 2025-11-09: -0 linesWeek of 2025-11-16: +0 linesWeek of 2025-11-16: -0 linesWeek of 2025-11-23: +0 linesWeek of 2025-11-23: -0 linesWeek of 2025-11-30: +0 linesWeek of 2025-11-30: -0 linesWeek of 2025-12-07: +0 linesWeek of 2025-12-07: -0 linesWeek of 2025-12-14: +0 linesWeek of 2025-12-14: -0 linesWeek of 2025-12-21: +0 linesWeek of 2025-12-21: -0 linesWeek of 2025-12-28: +0 linesWeek of 2025-12-28: -0 linesWeek of 2026-01-04: +0 linesWeek of 2026-01-04: -0 linesWeek of 2026-01-11: +0 linesWeek of 2026-01-11: -0 linesWeek of 2026-01-18: +0 linesWeek of 2026-01-18: -0 linesWeek of 2026-01-25: +0 linesWeek of 2026-01-25: -0 linesWeek of 2026-02-01: +2,646 linesWeek of 2026-02-01: -442 linesWeek of 2026-02-08: +1,154 linesWeek of 2026-02-08: -65 linesWeek of 2026-02-15: +1,389 linesWeek of 2026-02-15: -188 linesWeek of 2026-02-22: +1,001 linesWeek of 2026-02-22: -846 linesWeek of 2026-03-01: +691 linesWeek of 2026-03-01: -1,032 linesWeek of 2026-03-08: +596 linesWeek of 2026-03-08: -566 linesWeek of 2026-03-15: +382 linesWeek of 2026-03-15: -48 linesWeek of 2026-03-22: +194 linesWeek of 2026-03-22: -55 linesWeek of 2026-03-29: +0 linesWeek of 2026-03-29: -0 linesWeek of 2026-04-05: +7,466 linesWeek of 2026-04-05: -974 linesWeek of 2026-04-12: +2,380 linesWeek of 2026-04-12: -1,202 linesWeek of 2026-04-19: +14,391 linesWeek of 2026-04-19: -12,898 linesWeek of 2026-04-26: +2,872 linesWeek of 2026-04-26: -627 linesWeek of 2026-05-03: +5,545 linesWeek of 2026-05-03: -2,026 linesWeek of 2026-05-10: +0 linesWeek of 2026-05-10: -0 linesWeek of 2026-05-17: +1,074 linesWeek of 2026-05-17: -157 linesWeek of 2026-05-24: +2,669 linesWeek of 2026-05-24: -643 linesWeek of 2026-05-31: +793 linesWeek of 2026-05-31: -55 linesWeek of 2026-06-07: +1,885 linesWeek of 2026-06-07: -760 linesWeek of 2026-06-14: +1,093 linesWeek of 2026-06-14: -155 linesWeek of 2026-06-21: +1,635 linesWeek of 2026-06-21: -510 linesWeek of 2026-06-28: +1,079 linesWeek of 2026-06-28: -653 linesWeek of 2026-07-05: +8,255 linesWeek of 2026-07-05: -13 linesWeek of 2026-07-12: +1,073 linesWeek of 2026-07-12: -580 linesWeek of 2026-07-19: +2,634 linesWeek of 2026-07-19: -201 linesWeek of 2026-07-26: +5,167 linesWeek of 2026-07-26: -863 linesWeek of 2026-08-02: +4,474 linesWeek of 2026-08-02: -3,456 linesAug 10, 2025Aug 2, 2026
+72.5K lines added, -29K removed over the last year.

Commits per week

last 52 weeks
520Week of 2025-08-10: 0 commitsWeek of 2025-08-17: 0 commitsWeek of 2025-08-24: 0 commitsWeek of 2025-08-31: 0 commitsWeek of 2025-09-07: 0 commitsWeek of 2025-09-14: 0 commitsWeek of 2025-09-21: 0 commitsWeek of 2025-09-28: 0 commitsWeek of 2025-10-05: 0 commitsWeek of 2025-10-12: 0 commitsWeek of 2025-10-19: 0 commitsWeek of 2025-10-26: 0 commitsWeek of 2025-11-02: 0 commitsWeek of 2025-11-09: 0 commitsWeek of 2025-11-16: 0 commitsWeek of 2025-11-23: 0 commitsWeek of 2025-11-30: 0 commitsWeek of 2025-12-07: 0 commitsWeek of 2025-12-14: 0 commitsWeek of 2025-12-21: 0 commitsWeek of 2025-12-28: 0 commitsWeek of 2026-01-04: 0 commitsWeek of 2026-01-11: 0 commitsWeek of 2026-01-18: 0 commitsWeek of 2026-01-25: 0 commitsWeek of 2026-02-01: 6 commitsWeek of 2026-02-08: 10 commitsWeek of 2026-02-15: 21 commitsWeek of 2026-02-22: 35 commitsWeek of 2026-03-01: 22 commitsWeek of 2026-03-08: 25 commitsWeek of 2026-03-15: 4 commitsWeek of 2026-03-22: 4 commitsWeek of 2026-03-29: 0 commitsWeek of 2026-04-05: 5 commitsWeek of 2026-04-12: 20 commitsWeek of 2026-04-19: 9 commitsWeek of 2026-04-26: 16 commitsWeek of 2026-05-03: 52 commitsWeek of 2026-05-10: 0 commitsWeek of 2026-05-17: 1 commitsWeek of 2026-05-24: 22 commitsWeek of 2026-05-31: 5 commitsWeek of 2026-06-07: 20 commitsWeek of 2026-06-14: 1 commitsWeek of 2026-06-21: 3 commitsWeek of 2026-06-28: 14 commitsWeek of 2026-07-05: 2 commitsWeek of 2026-07-12: 15 commitsWeek of 2026-07-19: 1 commitsWeek of 2026-07-26: 28 commitsWeek of 2026-08-02: 40 commitsAug 10, 2025Aug 2, 2026
381 commits in the last 52 weeks.

When work happens

weekday and hour
SunMonTueWedThuFriSat036912151821Sun 0:00 — 0 commitsSun 1:00 — 0 commitsSun 2:00 — 1 commitsSun 3:00 — 1 commitsSun 4:00 — 0 commitsSun 5:00 — 0 commitsSun 6:00 — 0 commitsSun 7:00 — 0 commitsSun 8:00 — 0 commitsSun 9:00 — 0 commitsSun 10:00 — 0 commitsSun 11:00 — 0 commitsSun 12:00 — 2 commitsSun 13:00 — 0 commitsSun 14:00 — 1 commitsSun 15:00 — 0 commitsSun 16:00 — 0 commitsSun 17:00 — 1 commitsSun 18:00 — 0 commitsSun 19:00 — 0 commitsSun 20:00 — 0 commitsSun 21:00 — 0 commitsSun 22:00 — 0 commitsSun 23:00 — 1 commitsMon 0:00 — 0 commitsMon 1:00 — 0 commitsMon 2:00 — 0 commitsMon 3:00 — 0 commitsMon 4:00 — 0 commitsMon 5:00 — 0 commitsMon 6:00 — 0 commitsMon 7:00 — 1 commitsMon 8:00 — 0 commitsMon 9:00 — 0 commitsMon 10:00 — 5 commitsMon 11:00 — 8 commitsMon 12:00 — 5 commitsMon 13:00 — 5 commitsMon 14:00 — 14 commitsMon 15:00 — 22 commitsMon 16:00 — 8 commitsMon 17:00 — 6 commitsMon 18:00 — 2 commitsMon 19:00 — 1 commitsMon 20:00 — 0 commitsMon 21:00 — 7 commitsMon 22:00 — 9 commitsMon 23:00 — 6 commitsTue 0:00 — 0 commitsTue 1:00 — 0 commitsTue 2:00 — 0 commitsTue 3:00 — 0 commitsTue 4:00 — 0 commitsTue 5:00 — 0 commitsTue 6:00 — 0 commitsTue 7:00 — 3 commitsTue 8:00 — 2 commitsTue 9:00 — 2 commitsTue 10:00 — 4 commitsTue 11:00 — 18 commitsTue 12:00 — 3 commitsTue 13:00 — 6 commitsTue 14:00 — 14 commitsTue 15:00 — 11 commitsTue 16:00 — 8 commitsTue 17:00 — 14 commitsTue 18:00 — 2 commitsTue 19:00 — 5 commitsTue 20:00 — 5 commitsTue 21:00 — 5 commitsTue 22:00 — 12 commitsTue 23:00 — 3 commitsWed 0:00 — 0 commitsWed 1:00 — 0 commitsWed 2:00 — 0 commitsWed 3:00 — 0 commitsWed 4:00 — 0 commitsWed 5:00 — 0 commitsWed 6:00 — 0 commitsWed 7:00 — 0 commitsWed 8:00 — 3 commitsWed 9:00 — 1 commitsWed 10:00 — 4 commitsWed 11:00 — 7 commitsWed 12:00 — 6 commitsWed 13:00 — 8 commitsWed 14:00 — 16 commitsWed 15:00 — 19 commitsWed 16:00 — 16 commitsWed 17:00 — 16 commitsWed 18:00 — 5 commitsWed 19:00 — 3 commitsWed 20:00 — 1 commitsWed 21:00 — 1 commitsWed 22:00 — 2 commitsWed 23:00 — 1 commitsThu 0:00 — 2 commitsThu 1:00 — 0 commitsThu 2:00 — 1 commitsThu 3:00 — 0 commitsThu 4:00 — 0 commitsThu 5:00 — 0 commitsThu 6:00 — 0 commitsThu 7:00 — 0 commitsThu 8:00 — 1 commitsThu 9:00 — 5 commitsThu 10:00 — 21 commitsThu 11:00 — 12 commitsThu 12:00 — 7 commitsThu 13:00 — 10 commitsThu 14:00 — 4 commitsThu 15:00 — 7 commitsThu 16:00 — 5 commitsThu 17:00 — 14 commitsThu 18:00 — 6 commitsThu 19:00 — 3 commitsThu 20:00 — 2 commitsThu 21:00 — 0 commitsThu 22:00 — 10 commitsThu 23:00 — 4 commitsFri 0:00 — 0 commitsFri 1:00 — 0 commitsFri 2:00 — 0 commitsFri 3:00 — 0 commitsFri 4:00 — 0 commitsFri 5:00 — 0 commitsFri 6:00 — 0 commitsFri 7:00 — 0 commitsFri 8:00 — 0 commitsFri 9:00 — 7 commitsFri 10:00 — 4 commitsFri 11:00 — 11 commitsFri 12:00 — 15 commitsFri 13:00 — 6 commitsFri 14:00 — 7 commitsFri 15:00 — 11 commitsFri 16:00 — 14 commitsFri 17:00 — 4 commitsFri 18:00 — 6 commitsFri 19:00 — 1 commitsFri 20:00 — 1 commitsFri 21:00 — 1 commitsFri 22:00 — 0 commitsFri 23:00 — 9 commitsSat 0:00 — 1 commitsSat 1:00 — 0 commitsSat 2:00 — 0 commitsSat 3:00 — 0 commitsSat 4:00 — 0 commitsSat 5:00 — 0 commitsSat 6:00 — 0 commitsSat 7:00 — 0 commitsSat 8:00 — 0 commitsSat 9:00 — 0 commitsSat 10:00 — 0 commitsSat 11:00 — 1 commitsSat 12:00 — 1 commitsSat 13:00 — 0 commitsSat 14:00 — 1 commitsSat 15:00 — 0 commitsSat 16:00 — 1 commitsSat 17:00 — 0 commitsSat 18:00 — 0 commitsSat 19:00 — 0 commitsSat 20:00 — 0 commitsSat 21:00 — 0 commitsSat 22:00 — 0 commitsSat 23:00 — 0 commits
Commit volume by weekday and hour (UTC). Larger dots mean more commits.
DateListRankStars gained
Aug 7, 2026monthly#6+6,090
Aug 6, 2026monthly#6+5,874
Aug 5, 2026monthly#7+5,734
Aug 4, 2026monthly#8+5,579
Aug 3, 2026daily#7+442
Aug 3, 2026monthly#10+5,497
Aug 2, 2026monthly#9+4,740
Aug 2, 2026daily#7+442
Aug 1, 2026daily#1+628
Aug 1, 2026monthly#9+4,740
Jul 31, 2026daily#1+628
Jul 30, 2026daily#2+827
Jul 29, 2026daily#5+227
  • public-apis/public-apis

    A collective list of free APIs

    454.9K stars · Python

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    385.5K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • openclaw/openclaw

    Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞

    384.4K stars · TypeScript

  • donnemartin/system-design-primer

    Learn how to design large-scale systems. Prep for the system design interview. Includes Anki flashcards.

    362.2K stars · Python