iyeque-audio-processingAudio ingestion, analysis, transformation, and generation (Transcribe, TTS, VAD, Features).
Install via ClawdBot CLI:
clawdbot install iyeque/iyeque-audio-processingGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated Mar 21, 2026
Automatically transcribe podcast episodes using Whisper and extract audio features like sentiment or topic markers for content indexing. This enables searchable transcripts and metadata generation for podcast platforms.
Transcribe customer service calls in real-time and apply voice activity detection to segment conversations. This helps in analyzing call quality, identifying key issues, and generating summaries for training and compliance.
Convert lecture notes into audio via text-to-speech for accessibility and use feature extraction to analyze student engagement from recorded sessions. This supports e-learning platforms in creating inclusive and interactive materials.
Transcribe doctor-patient voice notes and extract features like speech patterns for preliminary diagnostics. This aids in documentation efficiency and supports telemedicine applications with automated insights.
Offer audio processing as a cloud-based service with tiered plans based on usage (e.g., transcription hours, TTS characters). Target small businesses and developers needing scalable audio tools without infrastructure setup.
License the skill's tools via API to enterprises for integration into their internal systems, such as call centers or content management platforms. Charge per API call or through enterprise contracts.
Provide bespoke solutions by customizing the skill for specific client needs, like adding proprietary models or integrating with existing workflows. Focus on industries with unique audio processing requirements.
💬 Integration Tip
Ensure ffmpeg and Python dependencies are installed via the provided commands, and use virtual environments like uv to manage package conflicts during deployment.
Scored Aug 20, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...