ACE-StepACE-Step 1.5Music model for structured songs with style tags, lyrics, BPM, key, language, and long durations.
ACE-StepACE-Step 1.5Music model for structured songs with style tags, lyrics, BPM, key, language, and long durations.
Massively multilingual zero-shot TTS with Voice Clone and Voice Design modes, prompt preprocessing, speed, duration, and diffusion controls.
Stability AIStable Audio 3 MediumFast high-quality music and sound generation with audio-to-audio editing, inpainting, and continuation.
MMAudioMMAudio V2 (Video to Audio)Video-to-audio model that adds synchronized ambience, Foley, and effects from video plus prompt.
Sony video/audio model for fast text sound effects or synchronized audio from video.
Instant CPU-efficient English narration with one consistent synthetic male voice.
Small English TTS model with many voices and natural American, British, and international accents.

Text-to-speech with preset premium timbres and precise style control

Text-to-speech with voice creation from natural language descriptions

Inworld AIInworld TTS-1.5 MiniLow-latency expressive text-to-speech optimized for real-time apps

Expressive text-to-speech with five voices, speech tags, and multilingual support
Reliable OpenAI text-to-speech for clear narration, dialogue, and assistant voice output.
Song model for natural vocals and rich arrangements with English or Chinese structured lyrics.
S1 is a multilingual text-to-speech model from Fish Audio. It is suited for voice applications that need broad emotional expression, using parenthetical controls to guide speaking style across its supported...
S2 Pro is a multilingual text-to-speech model from Fish Audio. It is suited for expressive narration and multi-speaker dialogue, with natural-language controls for speaking style and emotion.
MAI-Voice-2-Flash is a low-latency text-to-speech model from Microsoft for voice agents, assistants, call centers, accessibility, narration, and other interactive applications. It generates expressive 24 kHz mono speech across 15 languages...
Qwen-Audio-3.0-TTS Flash is Alibaba's fast, cost-efficient text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.
Qwen-Audio-3.0-TTS Plus is Alibaba's higher-quality text-to-speech model, generating spoken audio from text via the DashScope Speech Synthesizer API.
Aura-2 is a multilingual text-to-speech model from Deepgram. It supports Deepgram’s canonical Aura-2 voice catalog for speech synthesis across multiple languages.
MiniMax Speech 2.8 HD is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.
MiniMax Speech 2.8 Turbo is a text-to-speech model from MiniMax. It is suited for applications that generate spoken audio from text and accepts arbitrary MiniMax voice IDs.
MAI-Voice-2 is an expressive text-to-speech model from Microsoft. It is suited for conversational assistants, media narration, accessibility, education, and other long-form voice applications. It supports 15 languages across 18 locales,...

Fish AudioFish Audio S2.1 ProFlagship multilingual text-to-speech with natural language voice control and realtime streaming

Inworld AIInworld Realtime TTS-2Conversational text-to-speech with realtime voice direction and audio-aware delivery