qwen-audioHigh-performance audio library with text-to-speech (TTS) and speech-to-text (STT).
Install via ClawdBot CLI:
clawdbot install darknoah/qwen-audioGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen3-ASR-Repo/asr_en.wavAudited Apr 17, 2026 · audit v1.0
Generated Mar 21, 2026
Enables rapid creation of audiobooks or podcast episodes using high-quality TTS with customizable voices. Publishers can clone narrator voices for consistency across series and generate transcripts via STT for accessibility or marketing materials.
Integrates TTS to provide natural-sounding voice responses in IVR systems or chatbots, with voice cloning for brand-specific tones. STT transcribes customer calls for analysis, improving service quality and compliance.
Facilitates the development of interactive e-learning content by converting text lessons into speech with engaging, instructor-like voices. STT can transcribe student audio submissions for feedback or assessment.
Supports game developers and animators in generating character dialogues quickly using TTS with emotion control via instruct parameters. Voice cloning allows for unique character voices without hiring multiple actors.
Assists healthcare providers by transcribing patient consultations via STT for accurate medical records. TTS can convert written instructions into speech for patients with visual impairments, using calm, professional voices.
Offer the TTS and STT capabilities as cloud-based APIs with tiered pricing based on usage volume. Target businesses needing scalable audio processing, such as call centers or content creators, with premium features like advanced voice cloning.
License the skill as a customizable, on-premise solution for large organizations in industries like finance or healthcare. Provide integration support and maintenance contracts, ensuring data privacy and compliance with industry regulations.
Develop a user-friendly web or desktop application with free basic TTS/STT features and paid upgrades for high-quality voices, batch processing, or commercial use. Monetize through in-app purchases or premium subscriptions.
💬 Integration Tip
Ensure Python 3.10+ is installed and environment checks are completed before deployment; use the voice pre-check workflow to manage voice profiles efficiently.
Scored Apr 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...