Perso AI is a 3-in-1 AI audio and video platform combining AI Dubbing, Speech-to-Text, and Audio Separation in a single workflow. It translates and dubs videos into 33+ languages with natural voice cloning and lip dubbing, generates speaker-separated transcripts with automatic speaker diarization in four output formats (XLSX, SRT, VTT, JSON), and isolates individual speaker voices from background audio with dual modes (vocals-only or with reactions preserved). Trusted by 460,000+ users across 80+ countries, powered by the ElevenLabs voice engine (2025 partnership), and ISO/IEC 27001 and KISA ISMS certified. Developed by ESTsoft (est. 1993, KOSDAQ: 047560). Starts at $6.99/month with up to 98% cost savings vs. traditional dubbing studios.
AI Dubbing in 33+ languages, Speech-to-Text with speaker diarization, and Audio Separation — all in one workflow.
- Corporate L&D Teams — Translate training and onboarding videos. Speech-to-Text auto-generates meeting transcripts with speaker diarization.
- YouTube Creators — Dub videos into 33+ languages without hiring voice actors. Reach global audiences with localized content.
- Marketing Agencies — Localize campaign and product-demo videos for international markets while maintaining brand voice.
- Podcast Producers — Use Audio Separation to extract individual speaker voices, remove background music, or merge selected tracks for post-production.
- E-Learning Platforms — Translate lecture videos into regional languages. Auto-generate SRT subtitles for accessibility.
- Media Production — Create dubbed versions of documentary content for international distribution at a fraction of traditional dubbing costs.
- Sign up at perso.ai and upload any video or audio file. Choose your workflow — AI Dubbing to translate videos into 33+ languages
- Lip Dubbing for natural lip-synced output
- Speech-to-Text for speaker-separated transcripts and subtitles (XLSX
- SRT
- VTT
- JSON)
- or Audio Separation to isolate individual speakers and background audio. Edit the auto-generated script in the real-time editor for instant regeneration
- then download or export. Enterprise users can access all capabilities via the API for batch processing.

AI tool to animate photos with speech and lifelike expressions.


AI video lipsync tool for real-time lipsync and seamless translation.


Free online AI tool for creating lifelike lip-synced videos easily.


AI lip-sync talking video generator for lifelike, infinite-length performances.


VFX studio delivering feature-film quality VFX for TV series with innovative technology.


AI tool for generating realistic lip-synced talking videos from audio and images.


AI-powered lip sync video generator that turns voice or text into realistic talking videos in seconds.


AI-powered video translation tool supporting 130+ languages with lip sync and voice cloning.


AI tool for perfectly synchronized lip movements in videos.


Google's AI tool for generating videos with synchronized audio.



AI tool for realistic talking avatars and speech animations.


AI studio for videos, images, lip-sync, face swap, and music.


Platform for creating, animating, and deploying AI-driven virtual interactive personalities.


AI-powered tool for professional lip sync animation and auto audio-video synchronization.


AI platform for creating talking avatar videos with voice cloning and lip-sync.


All-in-one AI video and image creation platform.


AI-powered video and audio translation with lip sync and voice cloning.




