iyeque-audio-processingPrivacy-first local audio toolkit for OpenClaw: transcribe voice notes, create timestamped segments, extract features, detect speech, and run deterministic FFmpeg transforms. Use remote TTS only when explicitly requested and network consent is provided.
Install via ClawdBot CLI:
clawdbot install iyeque/iyeque-audio-processingGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated Mar 21, 2026
Automatically transcribe podcast episodes using Whisper and extract audio features like sentiment or topic markers for content indexing. This enables searchable transcripts and metadata generation for podcast platforms.
Transcribe customer service calls in real-time and apply voice activity detection to segment conversations. This helps in analyzing call quality, identifying key issues, and generating summaries for training and compliance.
Convert lecture notes into audio via text-to-speech for accessibility and use feature extraction to analyze student engagement from recorded sessions. This supports e-learning platforms in creating inclusive and interactive materials.
Transcribe doctor-patient voice notes and extract features like speech patterns for preliminary diagnostics. This aids in documentation efficiency and supports telemedicine applications with automated insights.
Offer audio processing as a cloud-based service with tiered plans based on usage (e.g., transcription hours, TTS characters). Target small businesses and developers needing scalable audio tools without infrastructure setup.
License the skill's tools via API to enterprises for integration into their internal systems, such as call centers or content management platforms. Charge per API call or through enterprise contracts.
Provide bespoke solutions by customizing the skill for specific client needs, like adding proprietary models or integrating with existing workflows. Focus on industries with unique audio processing requirements.
💬 Integration Tip
Ensure ffmpeg and Python dependencies are installed via the provided commands, and use virtual environments like uv to manage package conflicts during deployment.
Scored Aug 20, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.