Instant voice cloning by MIT and MyShell. Audio foundation model.
-
Updated
Apr 19, 2025 - Python
Instant voice cloning by MIT and MyShell. Audio foundation model.
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Foundational model for human-like, expressive TTS
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
Daily tracking of awesome audio papers, including music generation, zero-shot tts, asr, audio generation
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling
AuK: An Open-Source Foundational Model for Speech Generation and Editing
Benchmark for voice cloning robustness, speaker privacy, and audio protection across 26 TTS and VC models.
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
Dockerized Voicecraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
Arabic-first expressive zero-shot speech synthesis — Audar-TTS-V1 (Flash + Turbo + Pro). Voice cloning, expression tags, benchmarks, technical report & inference.
ComfyUI nodes for KRAFTON Raon-OpenTTS — open-weight, open-data zero-shot voice cloning (F5-TTS-style CFM/DiT + HiFi-GAN, 16 kHz English).
🎙️ OmniVoice Thai API — Zero-shot Thai TTS with Web UI + REST API (Voice Cloning, Voice Design, Auto Voice)
Self-hosted zero-shot voice cloning: turn a 3-30s sample into a reusable voice profile and synthesize speech from any text.
Bilingual IndexTTS 1.5 API service with local WebUI, voice management, and queued generation / IndexTTS 1.5 双语 API 服务,集成本地 WebUI、音色管理与队列式生成
Production-grade fine-tuning & LoRA toolkit for Chatterbox-Flash zero-shot TTS models. Combines parallel block diffusion and FlashInfer acceleration with smart placeholder vocabulary extension supporting languages. Features offline feature preprocessing, Silero VAD silence trimming, zero-padding leakage prevention, and fast voice cloning for custom
[Interspeech 2026] DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching
[ACL 2025] OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching
Browser-first zero-shot text-to-speech and voice cloning for JavaScript/TypeScript. WebGPU + ONNX, no Python, no backend, no API keys.
To associate your repository with the zero-shot-tts topic, visit your repo's landing page and select "manage topics."