Audio & Voice
Text-to-speech, voice cloning, transcription, and music generation
18 tools
Text-to-speech, voice cloning, transcription, and music generation
18 tools
We use analytics cookies (PostHog) including session replay and heatmaps to improve AI Jungle. No tracking happens before you choose. or read our Cookie Policy.
SayVocal is a TikTok voice generator offering 300+ AI voices for text-to-speech voiceovers, enabling creators to generate natural voices without requiring signup.
Dograh is an open source alternative to VAPI that provides voice AI and telephony capabilities for building conversational agents and voice applications.
India-first voice AI platform for building multilingual AI phone and web agents supporting 11+ Indian languages with bring-your-own-LLM flexibility.
Private macOS dictation software with 110+ language support, local-first processing, and multilingual speech-to-text for India-focused workflows.
Leaping AI is a voice and SMS platform that automates multi-day outreach campaigns for home remodeling companies with compliance tracking.
Estera is an AI receptionist that handles inbound calls, WhatsApp messages, and outbound prospecting with multi-channel lead qualification and appointment booking.
On-device voice transcription, rewriting, translation, and agentic assistant for macOS entirely local with no cloud, no accounts, runs on Apple Silicon.
Lispr is a free voice-to-text app for macOS that transcribes speech into text across 99 languages with translation support, requiring no account or subscription.
Deploy real-time AI agents that talk, type, and take action with voice synthesis, multimodal capabilities, and enterprise integration for customer support and business automation.
Deepgram helps developers build conversational AI using enterprise voice AI APIs for speech-to-text, text-to-speech, and real-time voice agent processing.
Cartesia Sonic is a real-time text-to-speech API with ultra-low latency (90ms) generating natural, expressive voices with laughter and emotion controls across 40+ languages for AI voice agents.
Noiz AI is a text-to-speech platform that generates emotionally expressive voice output using emoji-based tone control. It enables users to create natural, nuanced voices for storytelling and messaging.
Beatoven.ai generates royalty-free background music and soundscapes for videos and podcasts using AI composition technology.
Play.ht generates natural-sounding AI voiceovers from text, offering multiple voice options for content creators, marketers, and developers.
Krisp uses AI to remove background noise and distractions from audio during calls and recordings, improving voice quality in virtual meetings.
Synthwave is an AI music creation platform that generates original compositions and soundtracks in various genres for creative projects.
Vocapia uses AI to transcribe and index audio and video content for search and accessibility, helping organizations make media searchable.
Udio is an AI music generation platform that enables creators to compose, customize, and generate original music tracks using artificial intelligence.