asr-skillsThis skill should be used when the user asks to "transcribe audio", "transcribe video", "convert speech to text", "generate subtitles", "create captions", "i...
Install via ClawdBot CLI:
clawdbot install lgwanai/asr-skillsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 11, 2026
Transcribe business meetings from audio or video recordings to capture discussions, decisions, and action items. Speaker diarization identifies who said what, making it easy to assign tasks.
Generate SRT or ASS subtitle files for video content such as lectures, webinars, or training materials. Supports styled subtitles with speaker labels for better accessibility.
Transcribe court depositions or legal proceedings with accurate timestamps and speaker identification. Local processing ensures confidentiality of sensitive legal data.
Convert doctor-patient conversations or medical dictations into structured text with speaker labels. Helps maintain accurate patient records without third-party cloud services.
Transcribe podcasts, interviews, or focus groups into various formats like JSON or Markdown for content repurposing, SEO, or analysis. Speaker diarization helps attribute quotes to specific speakers.
Offer a cloud-hosted version of the ASR skill as a subscription service where users upload media for transcription. Revenue from monthly or annual fees based on usage tiers (e.g., hours processed).
License the skill to enterprises for on-premises deployment, ensuring data privacy and compliance. Revenue from upfront licensing fees plus annual maintenance and support.
Expose the ASR functionality as an API with a pay-per-minute or pay-per-file pricing model. Developers integrate it into their applications, and you charge per minute of audio transcribed.
💬 Integration Tip
The skill supports CLI, Python API, and async execution; for web integration, wrap the Python API with a REST endpoint using Flask or FastAPI.
Scored May 11, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...