Trending repositories: text-to-speech

18 tracked repositories tagged with text-to-speech, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

18 of 18 repositories

  • harry0703/MoneyPrinterTurbo

    利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

    AI summary: An automated AI workflow tool that generates high-definition short videos from a single topic or keyword.

    128,421ai-mlPythonMIT
  • unslothai/unsloth

    Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.

    AI summary: A local UI and framework for efficiently training and running large language models.

    77,160ai-mlPythonApache-2.0
  • calesthio/OpenMontage

    World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.

    AI summary: An open-source, automated video editing tool that dynamically pieces together clips based on text or audio transcripts.

    63,046developer-toolsPythonAGPL-3.0
  • jamiepine/voicebox

    The open-source AI voice studio. Clone, dictate, create.

    AI summary: A local, open-source AI voice generation platform built for real-time speech and global dictation.

    56,144ai-mlTypeScriptMIT
  • microsoft/VibeVoice

    Open-Source Frontier Voice AI

    AI summary: Open-source frontier voice AI models for long-form speech recognition and synthesis.

    54,602ai-mlPythonMIT
  • debpalash/VoiceStudio

    VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

    AI summary: A local-first application for voice cloning, dubbing, dictation, and long-form audio generation across 646 languages.

    53,052ai-mlPythonAGPL-3.0
  • OpenBMB/VoxCPM

    VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

    AI summary: A 2B-parameter tokenizer-free speech model for 30-language synthesis, voice cloning, and zero-shot voice design.

    38,289ai-mlPythonApache-2.0
  • KittenML/KittenTTS

    Open-source State-of-the-art TTS model which runs on a CPU 😻

    AI summary: A state-of-the-art, CPU-optimized Text-to-Speech model that operates under 25MB on disk.

    15,494ai-mlPythonApache-2.0
  • supertone-oss-archive/supertonic

    Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.

    AI summary: A blazing-fast, 99M-parameter multilingual text-to-speech engine optimized for on-device inference.

    13,782ai-mlSwiftMIT
  • QwenLM/Qwen3-TTS

    Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.

    AI summary: A highly natural, multilingual Text-to-Speech model from the Qwen team, capable of zero-shot voice cloning.

    13,626ai-mlPythonApache-2.0
  • huggingface/speech-to-speech

    Build voice agents with open-source models

    AI summary: A framework for building entirely voice-driven AI applications with incredibly low latency.

    13,363ai-mlPythonApache-2.0
  • abus-aikorea/voice-pro

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    AI summary: Comprehensive Gradio WebUI for advanced audio processing, featuring zero-shot voice cloning, transcription, and dubbing.

    12,970ai-mlPythonGPL-3.0
  • elebumm/RedditVideoMakerBot

    Create Reddit Videos with just✨ one command ✨

    AI summary: A bot that automatically creates compilation videos from Reddit threads using text-to-speech and programmatic editing.

    12,540otherPythonGPL-3.0
  • kyutai-labs/pocket-tts

    A TTS that fits in your CPU (and pocket)

    AI summary: A highly efficient, pocket-sized text-to-speech model designed for rapid, low-latency audio generation.

    9,754ai-mlPythonMIT
  • Blaizzy/mlx-audio

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

    AI summary: A comprehensive audio generation framework optimized natively for Apple Silicon using MLX.

    7,978ai-mlPythonMIT
  • Osmantic/ODS

    ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.

    AI summary: An automated deployment system that wires together Ollama, Open WebUI, and workflow tools to create a private AI server.

    6,969ai-mlPythonApache-2.0
  • MisoLabsAI/MisoTTS

    Miso TTS is an 8 billion, highly emotive text-to-speech model

    AI summary: A state-of-the-art 8B parameter Text-to-Speech model generating high-fidelity audio.

    3,239ai-mlPythonOther
  • samuel-vitorino/sopro

    A lightweight text-to-speech model with zero-shot voice cloning

    AI summary: A lightweight 135M parameter text-to-speech model capable of ultra-fast inference and zero-shot voice cloning.

    959ai-mlPythonApache-2.0