Industrial-level zero-shot text-to-speech system with voice cloning from a single audio clip
Open source AI toolkit for building real-time voice agents with on-device speech processing
Unified C++ speech engine with 43 ASR backends and 48 TTS engines in one binary, zero Python dependencies
Build local voice agents with open-source AI models
Turn e-books into audiobooks with AI-powered speech synthesis
AI-powered video translation with automatic dubbing, transcription, and subtitles
Open-source speech and sound generation models for high-fidelity, expressive audio synthesis
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
Instant zero-shot voice cloning with granular style control, by MIT & MyShell
Microsoft's open-source frontier voice AI for speech recognition and text-to-speech