doubao-podcastUse when calling Doubao/ByteDance podcast TTS API to generate audio, parsing WebSocket binary frames, handling streaming audio chunks, extracting audio_url f...
Install via ClawdBot CLI:
clawdbot install mileszhang001-boom/doubao-podcastGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://...Audited Apr 17, 2026 · audit v1.0
Generated May 22, 2026
Convert long articles (2000-25000 words) into natural-sounding podcast audio using ByteDance's TTS. Real-time streaming or non-streaming audio delivery enables immediate playback or caching.
Transform educational material or lesson transcripts into podcast-style audio with multiple speaker voices for an engaging learning experience. Supports random speaker ordering to maintain variety.
Leverage the podcast API to generate audiobooks from text or articles, using different voices for different characters or sections. The streaming API allows progressive audio ingestion.
Automatically convert news articles or blog posts into daily podcast summaries with customizable audio configurations (speech rate, background music). Ideal for news aggregators.
Generate podcast audio from localized content (e.g., regional news, community blogs) using appropriate speaker voices for the target audience. The API's input_text mode handles short updates efficiently.
Offer a subscription-based API service where users pay per minute of generated audio or monthly tiers based on volume. Caching avoids redundant costs.
Integrate podcast generation into a content platform; serve ads within the generated audio and share revenue with content creators. Streaming allows real-time ad insertion.
Provide white-label podcast generation for enterprises (e.g., internal training, marketing). Charge a setup fee plus monthly retainer for maintenance and API usage.
💬 Integration Tip
Use a server-side proxy to set required headers since browser WebSocket cannot set custom headers; handle PodcastEnd event promptly to avoid timeout waiting for SessionFinished.
Scored May 22, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.