speech-isolateVocal isolation / background music removal on remote (FREE) L4 GPU. Trigger when user says: isolate vocals, remove background music, extract voice, 提取人声, 去除背...
Install via ClawdBot CLI:
clawdbot install speech2srt/speech-isolateGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 6, 2026
A podcast editor needs to remove background music and noise from recorded episodes to isolate clean vocals for publishing. This skill automates the process using GPU-accelerated Demucs, saving hours of manual editing.
A YouTuber extracts vocals from a video with background music to create a clean audio track for subtitles or remixes. The tool handles batch processing of multiple video files.
A music producer isolates vocal stems from existing songs to create remixes or samples. The high-quality output from Demucs htdemucs_ft enables professional-grade separation.
A legal transcription service removes background noise from courtroom recordings to enhance speech clarity for accurate transcriptions. The tool's de-noise capability improves transcription accuracy.
A company providing closed captioning for webinars extracts clean vocals from recordings to generate accurate captions for hearing-impaired viewers. The automation reduces turnaround time.
Offer a web interface where users can upload files for vocal isolation with a free tier (limited minutes) and paid subscriptions for higher usage. Revenue from monthly or annual subscriptions.
Package the skill as an API for integration into media editing software or platforms (e.g., video editing tools). Charge per API call or a flat monthly fee for developers.
Sell the entire pipeline as a branded solution to post-production houses or studios. Revenue from one-time setup fees plus ongoing support and maintenance contracts.
💬 Integration Tip
Ensure Modal CLI is configured with valid token_id and token_secret before running; use a GPU-enabled Modal environment (e.g., A10G) for faster inference.
Scored Jul 9, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...