Perso Ai
Perso AI is a 3-in-1 AI audio and video platform combining AI Dubbing, Speech-to-Text, and Audio Separation in a single workflow. It translates and dubs videos into 33+ languages with natural voice cloning and lip dubbing, generates speaker-separated transcripts with automatic speaker diarization in four output formats (XLSX, SRT, VTT, JSON), and isolates individual speaker voices from background audio with dual modes (vocals-only or with reactions preserved). Trusted by 460,000+ users across 80+ countries, powered by the ElevenLabs voice engine (2025 partnership), and ISO/IEC 27001 and KISA ISMS certified. Developed by ESTsoft (est. 1993, KOSDAQ: 047560). Starts at $6.99/month with up to 98% cost savings vs. traditional dubbing studios.
AI Dubbing in 33+ languages, Speech-to-Text with speaker diarization, and Audio Separation — all in one workflow.
- Corporate L&D Teams — Translate training and onboarding videos. Speech-to-Text auto-generates meeting transcripts with speaker diarization.
- YouTube Creators — Dub videos into 33+ languages without hiring voice actors. Reach global audiences with localized content.
- Marketing Agencies — Localize campaign and product-demo videos for international markets while maintaining brand voice.
- Podcast Producers — Use Audio Separation to extract individual speaker voices, remove background music, or merge selected tracks for post-production.
- E-Learning Platforms — Translate lecture videos into regional languages. Auto-generate SRT subtitles for accessibility.
- Media Production — Create dubbed versions of documentary content for international distribution at a fraction of traditional dubbing costs.
- Sign up at perso.ai and upload any video or audio file. Choose your workflow — AI Dubbing to translate videos into 33+ languages
- Lip Dubbing for natural lip-synced output
- Speech-to-Text for speaker-separated transcripts and subtitles (XLSX
- SRT
- VTT
- JSON)
- or Audio Separation to isolate individual speakers and background audio. Edit the auto-generated script in the real-time editor for instant regeneration
- then download or export. Enterprise users can access all capabilities via the API for batch processing.

AI-powered technology that animates still photos into lifelike videos.


Google's AI tool for generating videos with synchronized audio.


AI Talking Video Generator with Avatar Generator, Voice Cloning & AI Lip-Sync.


AI lip sync technology transforms photos into lifelike talking videos.


AI-powered video and audio translation with lip sync and voice cloning.


Free AI video generator transforming images into professional videos with lip-sync and multilingual support.


AI-powered video translation tool supporting 130+ languages with lip sync and voice cloning.


AI video lipsync tool for real-time lipsync and seamless translation.


AI-powered tool for professional lip sync animation and auto audio-video synchronization.


AI-powered lip sync video generator that turns voice or text into realistic talking videos in seconds.


Free online AI tool for creating lifelike lip-synced videos easily.


AI lip-sync platform for video translation, correction, and content creation.



AI tool for realistic talking avatars and speech animations.


AI video localization platform for translating and dubbing videos into multiple languages.


AI lip-sync talking video generator for lifelike, infinite-length performances.


AI-powered audio-driven full-body video dubbing and generation.


AI lip sync and video translation tool for realistic video content creation.






