Perso Ai
Perso AI is a 3-in-1 AI audio and video platform combining AI Dubbing, Speech-to-Text, and Audio Separation in a single workflow. It translates and dubs videos into 33+ languages with natural voice cloning and lip dubbing, generates speaker-separated transcripts with automatic speaker diarization in four output formats (XLSX, SRT, VTT, JSON), and isolates individual speaker voices from background audio with dual modes (vocals-only or with reactions preserved). Trusted by 460,000+ users across 80+ countries, powered by the ElevenLabs voice engine (2025 partnership), and ISO/IEC 27001 and KISA ISMS certified. Developed by ESTsoft (est. 1993, KOSDAQ: 047560). Starts at $6.99/month with up to 98% cost savings vs. traditional dubbing studios.
AI Dubbing in 33+ languages, Speech-to-Text with speaker diarization, and Audio Separation — all in one workflow.
- Corporate L&D Teams — Translate training and onboarding videos. Speech-to-Text auto-generates meeting transcripts with speaker diarization.
- YouTube Creators — Dub videos into 33+ languages without hiring voice actors. Reach global audiences with localized content.
- Marketing Agencies — Localize campaign and product-demo videos for international markets while maintaining brand voice.
- Podcast Producers — Use Audio Separation to extract individual speaker voices, remove background music, or merge selected tracks for post-production.
- E-Learning Platforms — Translate lecture videos into regional languages. Auto-generate SRT subtitles for accessibility.
- Media Production — Create dubbed versions of documentary content for international distribution at a fraction of traditional dubbing costs.
- Sign up at perso.ai and upload any video or audio file. Choose your workflow — AI Dubbing to translate videos into 33+ languages
- Lip Dubbing for natural lip-synced output
- Speech-to-Text for speaker-separated transcripts and subtitles (XLSX
- SRT
- VTT
- JSON)
- or Audio Separation to isolate individual speakers and background audio. Edit the auto-generated script in the real-time editor for instant regeneration
- then download or export. Enterprise users can access all capabilities via the API for batch processing.

AI platform for creating talking avatar videos with voice cloning and lip-sync.


AI Talking Video Generator with Avatar Generator, Voice Cloning & AI Lip-Sync.


Audio-driven AI tool for talking avatars with precise lip sync.


AI tool for generating realistic lip-synced talking videos from audio and images.


AI video lipsync tool for real-time lipsync and seamless translation.


AI studio for videos, images, lip-sync, face swap, and music.


Can’t say “sorry,” “I love you,” or “thank you”? Let your pet do it. Upload a pet photo, pick an emotion scene, and type one sentence—get a natural lip-synced 5-second vertical video in 1–2 minutes.Turn pet photos into lip-synced videos that say what’s hard to say—sorry, love, or thanks—in minutes.


AI-powered tool for professional lip sync animation and auto audio-video synchronization.


AI video lipsync tool for real-time lipsync and seamless translation.


AI-powered technology that animates still photos into lifelike videos.


Video personalization platform using AI to create custom videos at scale.


AI-powered lip sync video generator that turns voice or text into realistic talking videos in seconds.


All-in-one AI video and image creation platform.


AI video localization platform for translating and dubbing videos into multiple languages.


AI platform to generate, edit, and translate talking videos with prompts.


AI lip sync and video translation tool for realistic video content creation.


Flawless uses AI to revolutionize filmmaking with visual dubbing and AI reshoots.


AI-powered lip-syncing for videos with text-to-speech in 90+ languages.




