doubao-asrTranscribe recorded audio files to text via Doubao Seed-ASR 2.0 (豆包录音文件识别模型2.0) from ByteDance/Volcengine. Best-in-class Chinese speech recognition with spea...
Install via ClawdBot CLI:
clawdbot install vahnxu/doubao-asrGrade Good — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://www.volcengine.com/docs/6561/1354868Audited Apr 17, 2026 · audit v1.0
Generated Mar 20, 2026
Transcribes recorded meetings from audio files (e.g., m4a, mp3) into text with speaker identification, enabling teams to review discussions and assign action items. Ideal for corporate environments where tracking who said what is crucial for follow-ups and documentation.
Converts voice memos or recorded interviews into text for journalists, researchers, or students to analyze content and extract quotes. Supports various audio formats like wav and flac, making it useful for field recordings or personal notes.
Transcribes customer service calls to text with speaker separation, helping companies monitor interactions, identify common issues, and train staff. Enhances quality assurance by providing searchable transcripts for compliance and improvement.
Transcribes audio recordings of legal proceedings or medical consultations into accurate text records, aiding in documentation and case management. Ensures precise transcripts for archival and reference purposes in regulated industries.
Converts podcast or video audio tracks into text for subtitles, show notes, or content repurposing. Streamlines production workflows by providing editable transcripts that can be used for SEO and audience engagement.
Offers monthly or annual plans for businesses to transcribe a set number of audio files, with tiered pricing based on usage volume. Generates recurring revenue by catering to regular needs like meeting recordings and customer calls.
Provides API access for developers to integrate transcription into their apps, charging per minute of audio processed. Attracts tech companies and startups needing scalable, on-demand speech recognition without upfront costs.
Sells customized licenses to corporations for unlimited transcription within their infrastructure, including support and compliance features. Targets industries like legal or healthcare with high-volume, sensitive audio processing needs.
💬 Integration Tip
Ensure all required environment variables are set correctly, especially the API key and TOS bucket details, to avoid upload and authentication errors during transcription.
Scored Jun 19, 2026
Use CodexBar CLI local cost usage to summarize per-model usage for Codex or Claude, including the current (most recent) model or a full model breakdown. Trigger when asked for model-level usage/cost data from codexbar, or when you need a scriptable per-model summary from codexbar cost JSON.
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
让 AI 代理根据对话内容自动选择最合适的模型。四层识别(系统过滤→关键词→指示词→语义相似度),四池架构(高速/智能/人文/代理),五分支路由,全自动 Fallback 回路。支持 trigger_groups_all 非连续词组命中。