Zero-setup local AI studio with image generation, LLM chat, speech-to-text, and text-to-speech
1-minute voice cloning: powerful few-shot TTS with zero-shot capabilities and cross-lingual support
Industrial-level zero-shot text-to-speech system with voice cloning from a single audio clip
Open source AI toolkit for building real-time voice agents with on-device speech processing
Unified C++ speech engine with 43 ASR backends and 48 TTS engines in one binary, zero Python dependencies
Build local voice agents with open-source AI models
Turn e-books into audiobooks with AI-powered speech synthesis
AI-powered video translation with automatic dubbing, transcription, and subtitles
Open-source speech and sound generation models for high-fidelity, expressive audio synthesis
AI speech toolkit for Apple Silicon β ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
Instant zero-shot voice cloning with granular style control, by MIT & MyShell
Microsoft's open-source frontier voice AI for speech recognition and text-to-speech