modesty-audiopodUse SkillBoss API Hub for audio processing tasks including AI music generation (text-to-music, instrumentals, samples), text-to-speech, speech-to-text transc...
Install via ClawdBot CLI:
clawdbot install modestyrichards/modesty-audiopodGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
POST → https://api.heybossai.com/v1/pilotCalls external URL not in known-safe list
https://skillboss.co/skill.mdAudited Apr 17, 2026 · audit v1.0
Generated Oct 6, 2026
Podcast producers use the skill to transcribe episodes, remove background noise, and separate speaker tracks for cleaner editing. They can also generate custom intro/outro music and instrumental beds from text prompts without licensing hassles.
Creators extract vocals and instrumentals from songs to produce karaoke tracks or backing audio for cover videos. They can then generate custom original instrumentals and samples for their channel branding using text prompts.
Game studios generate looping ambient soundscapes, drum loops, and mood-based samples for levels and menus. They also use noise reduction and stem separation to clean up recorded voice-over audio for in-game dialogue and trailers.
E-learning platforms convert course text into natural-sounding speech for narrated lessons and accessibility compliance. They use transcription to create searchable course captions and subtitle files from recorded video lectures.
Bedroom producers split commercial tracks into up to 16 stems for remixing, sampling, and forensic analysis. They use audio-to-audio style transfer and generated samples to sketch beats before committing to a full production.
Embed the SkillBoss audio pipeline inside a SaaS product and meter each transcription, generation, or stem-separation call against a credit balance. Customers avoid managing multiple vendor accounts while you capture margin on every API request.
Package the capabilities into a branded portal that marketing and podcast agencies use for client deliverables like TTS ads, noise-cleaned recordings, and custom jingles. Charge a flat monthly platform fee plus per-project fees for high-volume rendering.
Offer free limited monthly minutes of transcription and short music generations to attract creators, then monetize longer durations, lossless stems, and faster queue priority. Upsell pro plans to podcasters and indie musicians who exceed free limits.
💬 Integration Tip
Abstract the pilot() call into a small wrapper that handles request timeouts, retries, and response shape validation so every audio feature shares consistent error handling. Always load SKILLBOSS_API_KEY from environment variables and never hardcode it in client-side code.
Scored Oct 6, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.