Trending repositories: text-to-speech
12 tracked repositories tagged with text-to-speech, ordered by stars. Use the topic filters below to narrow further.
12 of 12 repositories
harry0703/MoneyPrinterTurbo
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
AI summary: A one-stop AI tool for automatically generating short videos from keywords.
102,072ai-mlPythonMITunslothai/unsloth
Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek-V4, GLM and other models.
AI summary: A local UI and framework for efficiently training and running large language models.
69,627ai-mlPythonApache-2.0calesthio/OpenMontage
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
AI summary: An open-source, automated video editing tool that dynamically pieces together clips based on text or audio transcripts.
45,695developer-toolsPythonAGPL-3.0OpenBMB/VoxCPM
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
AI summary: A highly efficient, tokenizer-free Text-to-Speech (TTS) generation model.
35,013ai-mlPythonApache-2.0KittenML/KittenTTS
State-of-the-art TTS model under 25MB 😻
AI summary: A fast, lightweight Text-to-Speech engine optimized for low-latency streaming applications.
15,292ai-mlPythonApache-2.0supertone-inc/supertonic
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
AI summary: A blazing-fast, 99M-parameter multilingual text-to-speech engine optimized for on-device inference.
13,606ai-mlSwiftMITQwenLM/Qwen3-TTS
Qwen3-TTS is an open-source series of TTS models developed by the Qwen team at Alibaba Cloud, supporting stable, expressive, and streaming speech generation, free-form voice design, and vivid voice cloning.
AI summary: A highly natural, multilingual Text-to-Speech model from the Qwen team, capable of zero-shot voice cloning.
12,836ai-mlPythonApache-2.0abus-aikorea/voice-pro
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
AI summary: Comprehensive Gradio WebUI for advanced audio processing, featuring zero-shot voice cloning, transcription, and dubbing.
12,072ai-mlPythonGPL-3.0kyutai-labs/pocket-tts
A TTS that fits in your CPU (and pocket)
AI summary: A lightweight, 100M parameter text-to-speech model optimized to run efficiently on CPUs.
8,070ai-mlPythonMITBlaizzy/mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
AI summary: High-performance audio processing library optimized for Apple Silicon via the MLX framework.
7,688ai-mlPythonMITMisoLabsAI/MisoTTS
Miso TTS is an 8 billion, highly emotive text-to-speech model
AI summary: A state-of-the-art 8B parameter text-to-speech RVQ Transformer model.
3,192ai-mlPythonOthersamuel-vitorino/sopro
A lightweight text-to-speech model with zero-shot voice cloning
AI summary: A lightweight, streaming text-to-speech model offering zero-shot voice cloning at extreme speeds on CPU.
877ai-mlPythonApache-2.0