Miso One
Miso One is an open-weights, 8B-parameter text-to-speech (TTS) system developed by Miso Labs. It is designed specifically for producing highly realistic, expressive, and emotionally varied English conversational speech, making it ideal for voice-agent research and developer workflows. Built on a Sesame-style conversational speech model (CSM) architecture with Mimi audio codes, it features a highly optimized inference capability boasting a published low latency of 110 ms. In addition to text-to-speech generation, the model supports voice continuation and one-shot voice cloning from audio context with clear
Open-weights 8B AI text-to-speech model for expressive English speech.
- Running low-latency conversational speech generation locally for voice-agent research
- Generating high-quality, expressive voiceovers, narrations, and captions for creator content
- Creating instant voice clones and personal voice models using brief, consented audio context prompts
- Users can evaluate Miso One by reading its official model card on the repository or Hugging Face page
- trying out the hosted web demo to check voice quality
- or downloading the public 8B weights and inference code to run local benchmarks within their own CUDA environment. For hosted creator workflows
- users can sign up and choose a subscription plan based on their required annual or monthly character capacity.

AI voice generator with realistic text-to-speech and speech-to-speech capabilities.


A free web UI using OpenAI API to convert text to speech.


Text-to-speech Chrome extension for reading aloud digital content in multiple languages.


Free online text-to-speech converter supporting multiple languages.


Free online AI text to speech converter with natural voices and download options.


Free and open-source online AI voice generator for text to speech conversion.


A tool to download Microsoft synthesized Text-to-Speech audio with one click.


AI voice generator with 300 voices in 70+ languages for lifelike speech synthesis.


AI-powered multilingual voice synthesis and cloning platform with natural language processing.





