Trending repositories: tts

12 tracked repositories tagged with tts, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

12 of 12 repositories

  • unslothai/unsloth

    Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.

    AI summary: A local UI and framework for efficiently training and running large language models.

    69,627ai-mlPythonApache-2.0
  • microsoft/VibeVoice

    Open-Source Frontier Voice AI

    AI summary: An open-source framework for building voice-driven conversational agents.

    52,134ai-mlPythonMIT
  • OpenBMB/VoxCPM

    VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

    AI summary: A highly efficient, tokenizer-free Text-to-Speech (TTS) generation model.

    35,013ai-mlPythonApache-2.0
  • ATH-MaaS/Pixelle-Video

    🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine

    AI summary: AI Fully Automated Short Video Engine.

    26,529ai-mlPythonApache-2.0
  • ATH-MaaS/Pixelle-Video

    🚀 AI 全自动短视频引擎 | AI Fully Automated Short Video Engine

    AI summary: AI Fully Automated Short Video Engine by AIDC.

    26,157ai-mlPythonApache-2.0
  • KittenML/KittenTTS

    State-of-the-art TTS model under 25MB 😻

    AI summary: A fast, lightweight Text-to-Speech engine optimized for low-latency streaming applications.

    15,292ai-mlPythonApache-2.0
  • supertone-inc/supertonic

    Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

    AI summary: A blazing-fast, 99M-parameter multilingual text-to-speech engine optimized for on-device inference.

    13,606ai-mlSwiftMIT
  • QwenLM/Qwen3-TTS

    Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

    AI summary: A highly natural, multilingual Text-to-Speech model from the Qwen team, capable of zero-shot voice cloning.

    12,836ai-mlPythonApache-2.0
  • abus-aikorea/voice-pro

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    AI summary: Comprehensive Gradio WebUI for advanced audio processing, featuring zero-shot voice cloning, transcription, and dubbing.

    12,072ai-mlPythonGPL-3.0
  • moonshine-ai/moonshine

    Very low latency speech to text, intent recognition, and text to speech, for building voice agents and interfaces

    AI summary: A fast, cross-platform family of open speech-to-text models designed specifically for low-latency live voice interfaces.

    10,659ai-mlC++Other
  • GetStream/Vision-Agents

    Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.

    AI summary: Multi-modal AI agents by Stream for real-time video understanding with sub-30ms latency.

    8,006ai-mlPythonApache-2.0
  • Blaizzy/mlx-audio

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

    AI summary: High-performance audio processing library optimized for Apple Silicon via the MLX framework.

    7,688ai-mlPythonMIT