2026年6月28日

W

A general-purpose speech recognition model by OpenAI.
AI WritingContent GenerationResearchEmail WritingSummarizationRewritingAcademic Research浏览器扩展

概览
W 是什么?

Whisper is a general-purpose speech recognition model developed by OpenAI. It is trained on a large dataset of diverse audio and is also a multi-task model that can perform multilingual speech recognition as well as speech translation and language identification. Whisper uses a Transformer sequence-to-sequence model trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection. These tasks are jointly represented as a sequence of tokens to be predicted by the decoder, allowing a single model to replace many stages of a traditional speech-processing pipeline. The multitask training format uses a set of special tokens that serve as task specifiers or classification targets.

A general-purpose speech recognition model by OpenAI.

核心功能
Multilingual speech recognition
Speech translation
Language identification
Voice activity detection
热门使用场景
  • Transcribing audio files to text
  • Translating speech from one language to another
  • Identifying the language spoken in an audio file
如何使用
  • Whisper can be used via command-line or within Python. For command-line usage
  • you can transcribe speech in audio files by specifying the audio file and model size. For Python usage
  • you can load the model and use the transcribe() method to process audio files.
产品时间线
待核实
定价
W 采用 Free 定价模式,价格和功能可能会随时间变化。
Free
$0
待核实
Pro
待核实
待核实
Team
待核实
待核实
Enterprise
待核实
待核实
优惠 / 优惠码
暂无优惠码。
验证信息
工具状态
待核实
定价已核验
待核实
创始人已认领
否 / 待核实
来源
官网 / 社区提交
相关标签
AI WritingContent GenerationResearchEmail WritingSummarizationRewritingAcademic Research浏览器扩展Freemium
你是这个工具的官方团队吗?
认领这个资料页后,你可以更新产品信息、定价和官方回复。