openclaw-whisper-voiceLocal Whisper speech-to-text for audio files and inbound voice notes on the OpenClaw Gateway host. Use when setting up local transcription for WhatsApp, Tele...
Install via ClawdBot CLI:
clawdbot install sabyaghosh/openclaw-whisper-voiceGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://bootstrap.pypa.io/get-pip.pyAudited Apr 16, 2026 · audit v1.0
Generated Oct 6, 2026
A support team receives WhatsApp and Telegram voice notes from customers and needs them transcribed instantly on the Gateway host. This skill installs Whisper and ffmpeg locally and patches tools.media.audio to transcribe each inbound note automatically. Agents can then search, tag, and respond to transcripts without any cloud API dependency.
Technicians in the field send audio reports via messaging apps where network bandwidth is limited and privacy is critical. The skill provides a reliable transcribe.sh wrapper that converts .ogg, .m4a, or .mp3 files into text locally on the Gateway host. Reports are indexed and routed to dispatch systems without uploading audio to third-party services.
Clinics collect patient voice messages through secure chat channels and must keep all data on-premises to meet HIPAA-style requirements. This skill enables local transcription using base or small Whisper models, avoiding cloud APIs entirely. Transcripts can be reviewed by staff and attached to patient records without external exposure.
Independent podcasters and video editors receive raw interview audio and need quick transcripts for show notes, subtitles, and searchable archives. The skill supports standalone file transcription with --format srt, vtt, or txt for post-production tools. It replaces manual transcription or recurring cloud subscription costs with a one-time local setup.
Community managers on WhatsApp and Telegram groups need to understand voice notes in multiple languages for moderation and member support. The skill's translate task and model options let operators convert speech to English text or their own language on the host. This keeps moderation fast, private, and independent of per-minute cloud pricing.
Agencies and IT teams deploy the OpenClaw Gateway host with this skill to offer private, unlimited transcription for internal communication channels. They avoid cloud API usage fees and data residency concerns. The value proposition is predictable infrastructure cost plus full data control.
A B2B SaaS platform for messaging automation bundles this skill as an on-premise transcription module for enterprise customers. Customers get local Whisper processing while the vendor maintains the wrapper scripts and OpenClaw integration. This differentiates the platform against competitors relying on cloud speech APIs.
A solutions provider builds turnkey integrations for healthcare, legal, or field service clients using this skill as the transcription engine. They deliver configured Gateway hosts, custom routing, and transcript workflows tailored to each vertical. The skill reduces engineering time and avoids ongoing API costs.
💬 Integration Tip
Run scripts/install_local_whisper.sh first to create stable whisper and ffmpeg launchers under ~/.local/bin. Patch tools.media.audio with the transcribe.sh CLI entrypoint using --stdout-only so inbound voice notes return clean transcript text.
Scored Oct 6, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.