voice-stt-ttsFull voice message setup (STT + TTS) for OpenClaw using faster-whisper and Edge TTS
Install via ClawdBot CLI:
clawdbot install aksenkin/voice-stt-ttsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://docs.openclaw.ai/nodes/audioAudited Apr 16, 2026 · audit v1.0
Generated Mar 21, 2026
Integrate voice messaging into customer support systems to allow users to send voice queries and receive spoken responses. This reduces typing effort and improves accessibility for visually impaired users, enhancing overall customer experience.
Use STT for pronunciation practice and TTS for feedback in language learning apps. Students can record speech for transcription and correction, while the system provides spoken examples to aid comprehension and fluency development.
Enable healthcare professionals to record voice notes during patient consultations, which are transcribed for documentation. TTS can be used to read back notes or provide reminders, streamlining administrative tasks and reducing errors.
Transcribe podcast episodes automatically using STT for subtitles or written content, and use TTS to generate audio previews or summaries. This saves time on manual transcription and expands content reach through multiple formats.
Develop tools that convert spoken input to text for deaf or hard-of-hearing users and text to speech for blind users. This promotes inclusivity in digital communication, such as in messaging apps or workplace software.
Offer a cloud-based voice messaging API with pay-per-use or tiered subscription plans for businesses. Revenue comes from monthly fees based on usage volume, such as transcription minutes or TTS characters processed.
Provide basic STT and TTS features for free with limited usage, then charge for advanced features like higher accuracy models, multiple languages, or custom voice options. This attracts users and upsells premium services.
Sell customized voice messaging solutions to large enterprises with on-premise deployment, dedicated support, and integration services. Revenue is generated through one-time licensing fees and ongoing maintenance contracts.
💬 Integration Tip
Ensure the virtual environment is correctly set up and test the transcription script with sample audio files before full deployment to avoid runtime errors.
Scored Jun 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...