Miso One
Miso One is an open-weights, 8B-parameter text-to-speech (TTS) system developed by Miso Labs. It is designed specifically for producing highly realistic, expressive, and emotionally varied English conversational speech, making it ideal for voice-agent research and developer workflows. Built on a Sesame-style conversational speech model (CSM) architecture with Mimi audio codes, it features a highly optimized inference capability boasting a published low latency of 110 ms. In addition to text-to-speech generation, the model supports voice continuation and one-shot voice cloning from audio context with clear
Open-weights 8B AI text-to-speech model for expressive English speech.
- Running low-latency conversational speech generation locally for voice-agent research
- Generating high-quality, expressive voiceovers, narrations, and captions for creator content
- Creating instant voice clones and personal voice models using brief, consented audio context prompts
- Users can evaluate Miso One by reading its official model card on the repository or Hugging Face page
- trying out the hosted web demo to check voice quality
- or downloading the public 8B weights and inference code to run local benchmarks within their own CUDA environment. For hosted creator workflows
- users can sign up and choose a subscription plan based on their required annual or monthly character capacity.

AI voice solution for content creation with text-to-speech, dubbing, and voice cloning.


Free online AI text to speech generator with realistic voices and customization.


Free AI text-to-speech tool using ChatGPT voices for an immersive listening experience.

Converts live stream chat messages into voice.


A free web UI using OpenAI API to convert text to speech.


Guide and production API platform for advanced AI voice, speech, and music.


AI platform converts novels to audiobooks with unique character voices.


Web Whisper converts web pages into audio for podcast-like listening.


AI voice generator with realistic text-to-speech and speech-to-speech capabilities.







