xiaomi-any2speech-beyondtts声音世界模型(Speech World Model):不只是 TTS,而是理解场景、角色、情绪并自主规划表达的语音大模型。 原生支持长文+多人、中英双语,也支持上传参考音频进行音色克隆(Voice Prompt / voice cloning),内置高能创作模板,将任意内容转为播客/有声书/相声/Rap/广播剧等...
Install via ClawdBot CLI:
clawdbot install whiteshirt0429/xiaomi-any2speech-beyondttsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://miplus-tts-public.ai.xiaomi.comUses known external API (expected, informational)
open.feishu.cnAudited Apr 18, 2026 · audit v1.0
Generated Aug 5, 2026
用户提供文字内容(如文章、小说、博客)和参考音频,系统自动生成包含多角色、情感起伏的有声书或播客,支持中英双语。适用于内容创作者、自媒体人快速制作音频内容。
企业利用该技能构建智能语音应答系统,通过定制音色(如品牌专属声音)和多种风格(如客服、促销)提升客户体验。支持批量生成多语种语音提示、IVR 语音、产品介绍音频。
教师和课程开发者可生成多角色对话(如辩论、相声)用于语言学习,或制作带情感表达的课文朗读。学生可克隆自己的声音练习发音。
营销团队将文案转化为有吸引力的音频广告,如播客插播、电台广告、短视频配音。支持多种风格(Rap、脱口秀)和音色定制,提升品牌记忆点。
为视障用户提供高质量的文本到语音服务,支持自然情感表达和多角色叙事,改善阅读体验。同时可用于有声新闻、电子书朗读等。
面向开发者或企业提供API接口,按请求次数或处理时长收费。客户可直接集成到自己的应用中,享受专业级TTS和语音克隆功能。
建立在线平台,用户按月/年订阅即可使用高级功能(如无限制生成、多角色模板、商业授权)。提供免费试用吸引用户。
为大型企业提供定制化音色、场景模型和系统集成服务,包括培训、维护和专属支持。按项目或年度合同收费。
💬 Integration Tip
利用内置免费Key快速试用,注意处理异步任务和文件上传场景,可嵌入现有内容管理系统。
Scored Aug 5, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...