Trending repositories: speech-recognition

4 tracked repositories tagged with speech-recognition, ordered by stars. Use the topic filters below to narrow further.

Filter by topic

4 of 4 repositories

  • Vaibhavs10/insanely-fast-whisper

    AI summary: An optimized implementation of OpenAI's Whisper model that runs transcription incredibly fast.

    13,041ai-mlJupyter NotebookApache-2.0
  • abus-aikorea/voice-pro

    Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

    AI summary: Comprehensive Gradio WebUI for advanced audio processing, featuring zero-shot voice cloning, transcription, and dubbing.

    12,072ai-mlPythonGPL-3.0
  • Blaizzy/mlx-audio

    A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

    AI summary: High-performance audio processing library optimized for Apple Silicon via the MLX framework.

    7,688ai-mlPythonMIT
  • QwenLM/Qwen3-ASR

    Qwen3-ASR is an open-source series of ASR models developed by the Qwen team at Alibaba Cloud, supporting stable multilingual speech/music/song recognition, language detection and timestamp prediction.

    AI summary: An open-source suite of Automatic Speech Recognition models supporting multilingual transcription and timestamp prediction.

    3,316ai-mlPythonApache-2.0