video-audio-replaceReplace video audio with TTS voice while preserving original timing. Includes subtitle generation from video using Whisper. Uses ElevenLabs or Edge TTS, alig...
Install via ClawdBot CLI:
clawdbot install synthere/video-audio-replaceGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://api.elevenlabs.io/v1/text-to-speech/{voice_id}Audited Apr 16, 2026 · audit v1.0
Generated Mar 20, 2026
Content creators can dub their videos into multiple languages using TTS voices, expanding their audience reach without hiring voice actors. This is ideal for YouTube channels, educational content, or marketing videos where quick localization is needed. The tool preserves original timing, ensuring a natural viewing experience.
Educational institutions and e-learning platforms can use this skill to generate voice-overs for video lectures, making content accessible to visually impaired learners. By integrating Whisper for subtitle generation, it automates the creation of synchronized audio descriptions, improving compliance with accessibility standards.
Businesses can localize internal training videos for global teams by replacing audio with TTS in different languages, reducing production costs and time. The speed adjustment feature ensures the audio matches the original video pacing, maintaining instructional clarity across regions.
Social media managers can repurpose existing video content by replacing audio with trending TTS voices or different languages to engage diverse audiences. This allows for quick updates to viral videos or ads without re-recording, enhancing content strategy efficiency.
Offer a cloud-based platform where users upload videos and SRT files to generate TTS-dubbed versions via API. Charge monthly subscriptions based on usage tiers, such as video length or number of languages. This model targets businesses needing regular localization.
Provide a free version with basic Edge TTS and limited processing, then upsell to paid plans for ElevenLabs integration, advanced speed controls, and batch processing. Monetize through in-app purchases or annual licenses, appealing to individual creators and small teams.
License the skill to large enterprises, such as media companies or educational institutions, for integration into their internal workflows. Offer custom support, white-labeling, and volume discounts, generating revenue through long-term contracts and service fees.
💬 Integration Tip
Integrate this skill into existing video editing pipelines by automating SRT generation and TTS calls via API, ensuring compatibility with tools like FFmpeg for seamless audio replacement.
Scored Apr 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Search and manage Spotify playlists, tracks, albums, artists, and playback state via the Spotify Web API. Use this skill when users want to search for music,...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Transcribe long audio files safely on 16GB RAM machines using auto-chunking with Whisper’s base model and seamless transcript merging.
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...