sst-simpleLocal speech-to-text using OpenAI Whisper. Use when the user needs to: (1) transcribe audio files to text, (2) convert voice messages to written content, (3)...
Install via ClawdBot CLI:
clawdbot install lkisme/sst-simpleGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
rm -rf /Audited Apr 16, 2026 · audit v1.0
Generated Mar 21, 2026
Transcribe customer voice messages from platforms like WhatsApp or Telegram in multiple languages, enabling support teams to quickly understand and respond to inquiries. This is useful for global companies handling support tickets in languages such as Chinese, English, and Japanese.
Convert audio recordings of business meetings or conferences into text for documentation, note-taking, or creating subtitles. Supports batch processing of multiple files, ideal for organizations that need accurate records of discussions in various formats like .wav or .mp3.
Transcribe audio from interviews, podcasts, or videos to generate written content, subtitles in .srt or .vtt formats, or scripts for editing. Useful for media agencies and creators who require high-quality transcriptions in languages like Spanish or French.
Convert lecture recordings or training sessions into text for accessibility, study notes, or multilingual educational resources. Supports 99+ languages, making it suitable for online learning platforms and institutions with diverse student populations.
Offer a monthly or annual subscription for businesses to transcribe audio files with features like batch processing and multi-language support. Revenue is generated through tiered pricing based on usage volume and model accuracy levels.
Provide an API that developers can integrate into applications for on-demand speech-to-text conversion, charging per transcription minute or file. This model targets tech companies and app developers needing scalable, isolated outputs for multi-agent systems.
Sell enterprise licenses to large organizations requiring session isolation for multiple agents, such as customer support teams using WhatsApp and Telegram. Revenue comes from one-time or annual licensing fees with premium support and customization options.
💬 Integration Tip
Ensure the installation script is run first to set up the virtual environment and download models, and use session IDs for multi-agent scenarios to avoid output conflicts.
Scored Jun 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Turn your AI into JARVIS. Voice, wit, and personality — the complete package. Humor cranked to maximum.
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch contro...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...