Open Source
Explore the latest AI open-source projects from GitHub and HuggingFace.
Explore the latest AI open-source projects from GitHub and HuggingFace.
Index-TTS is an industrial-level, controllable zero-shot text-to-speech system that introduces precise duration control into an autoregressive TTS architecture while preserving natural prosody from voice prompts. It achieves disentanglement of speaker identity and emotional expression, enabling independent control over timbre and emotion through audio reference, emotional vector specification, or free-text description, with support for Chinese and English and optional FP16 and DeepSpeed acceleration.
RVC-Boss
Open-source WebUI for few-shot and zero-shot voice cloning and text-to-speech, producing a usable voice from as little as a 5-second sample.
Microsoft
Microsoft's MIT-licensed open frontier voice AI: 1.5B long-form TTS up to 90 minutes with 4 speakers, 0.5B streaming TTS at 300 ms latency, and 7B ASR for 60-minute single-pass transcription. 47k+ stars.
2noise
A dialogue-optimized open TTS model trained on 100,000+ hours that adds fine-grained prosody — laughter, pauses, interjections — with multi-speaker, English/Chinese support.
suno-ai
Suno's fully generative text-to-audio model — speech, music, and sound effects from one transformer, with nonverbal cues like [laughs] and 100+ voice presets (MIT).