moss-ttsdMOSI Studio 多人对话合成(moss-ttsd):将多个角色的对话文本合成为 单段连续音频,多人声音自然交替。支持 1~5 个说话人。 触发词:多说话人、多人对话、对话合成、多个角色、多种声音、多个人说话、 "multi-speaker"、"dialogue synthesis"、"多人对话"。 在飞书...
Install via ClawdBot CLI:
clawdbot install mkkb473/moss-ttsdGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://studio.mosi.cnAudited Apr 17, 2026 · audit v1.0
Generated Mar 21, 2026
Create interactive language learning or educational content where two characters engage in conversations to demonstrate dialogues. This helps learners practice listening comprehension in realistic scenarios, such as ordering food or asking for directions.
Generate audio dialogues between a customer and a service representative for training purposes. This allows employees to practice handling common inquiries or complaints in a controlled, auditory format.
Produce segments of audiobooks or podcasts where two characters converse, enhancing storytelling with distinct voices. This adds depth to narratives without requiring multiple voice actors.
Develop dialogues for voice assistants or chatbots that simulate natural conversations between two personas. This can be used in smart home devices or customer support bots to make interactions more engaging.
Create promotional audio content featuring two speakers, such as a testimonial dialogue or a product demonstration conversation. This makes ads more dynamic and relatable for listeners.
Offer the dialogue synthesis service through a monthly or annual subscription plan for developers and businesses. This provides predictable revenue and encourages long-term usage in applications like e-learning platforms.
Allow free limited usage for individual creators or small projects, with paid tiers for higher volume or advanced features. This attracts a broad user base while monetizing heavy users in industries like media production.
License the technology to large corporations for internal use, such as in training modules or customer service tools. This includes customization and integration support, targeting sectors like corporate training.
💬 Integration Tip
Ensure the MOSI_TTS_API_KEY environment variable is set and use the provided script with clear S1/S2 text formatting for seamless audio generation.
Scored Jun 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...