Miso One is an open-weights, 8B-parameter text-to-speech (TTS) system developed by Miso Labs. It is designed specifically for producing highly realistic, expressive, and emotionally varied English conversational speech, making it ideal for voice-agent research and developer workflows. Built on a Sesame-style conversational speech model (CSM) architecture with Mimi audio codes, it features a highly optimized inference capability boasting a published low latency of 110 ms. In addition to text-to-speech generation, the model supports voice continuation and one-shot voice cloning from audio context with clear
Open-weights 8B AI text-to-speech model for expressive English speech.
- Running low-latency conversational speech generation locally for voice-agent research
- Generating high-quality, expressive voiceovers, narrations, and captions for creator content
- Creating instant voice clones and personal voice models using brief, consented audio context prompts
- Users can evaluate Miso One by reading its official model card on the repository or Hugging Face page
- trying out the hosted web demo to check voice quality
- or downloading the public 8B weights and inference code to run local benchmarks within their own CUDA environment. For hosted creator workflows
- users can sign up and choose a subscription plan based on their required annual or monthly character capacity.
Azure AI Speech creates more human-like voices using improved AI speech synthesis.


AI-powered platform to transform text into engaging AI podcasts quickly and easily.


Online text-to-speech converter with natural voices and multiple formats support.


Free online AI-powered text-to-speech solution with multilingual support and voice cloning.


Free AI text-to-speech tool using ChatGPT voices for an immersive listening experience.


AI reading and listening assistant that converts text to audio and creates summaries.


Text-to-speech Chrome extension for reading aloud digital content in multiple languages.


Web Whisper converts web pages into audio for podcast-like listening.

Converts live stream chat messages into voice.








