multilingual-video-dubbing-text-to-speechPractical mastering steps for TTS audio: cleanup, loudness normalization, alignment, and delivery specs.
Install via ClawdBot CLI:
clawdbot install lnj22/multilingual-video-dubbing-text-to-speechGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated Sep 29, 2026
Creators generate AI voiceovers for documentary, explainer, or listicle videos using TTS. They need to clean up raw TTS output, normalize loudness to broadcast standards, and align each segment to video timing windows. This skill provides the exact mastering chain and sync tolerance to produce professional-sounding audio without a recording booth.
Training companies localize courses into multiple languages by generating TTS audio for each slide and stitching them into timed lessons. They must ensure consistent loudness across segments and avoid clicks or pops at boundaries. The skill's segment cleanup, padding, and re-normalization rules ensure a polished learning experience.
Podcast networks dynamically insert AI-generated ad reads into episodes. They need the ad audio to match the show's loudness target (-23 LUFS) and avoid jarring transitions. This skill guides measuring integrated loudness and applying final normalization after any duration edits.
Publishers convert manuscripts to audiobooks using TTS and need chapter-level consistency in tone and loudness. They must handle variable segment lengths, apply fades at chapter boundaries, and ensure true peaks stay below -1.5 dBTP. This skill offers practical FFmpeg patterns and timing guidelines.
Telephony teams update IVR menus with TTS prompts and require clean, consistent audio that meets telecom loudness specs. They clean up rumble and digital fizz, normalize to -23 LUFS, and align prompts to tight time windows. The skill's emphasis on native sample rate and boundary fades ensures callers hear clear, professional prompts.
A cloud-based service that automates TTS cleanup, loudness normalization, and segment alignment for video creators and podcasters. Users upload raw TTS files and receive broadcast-ready audio with a single click, using the skill's FFmpeg workflow under the hood. Subscription tiers are based on processing minutes and export formats.
An agency offers end-to-end TTS audio mastering as a white-label service for video production and e-learning companies. They handle engine selection, speech cleanup, loudness normalization, and timing alignment to meet client delivery specs. Pricing is per finished minute or per project.
A REST API that developers integrate into their own apps to automatically master TTS audio. It accepts raw audio and target specs, then returns cleaned and normalized files with segment boundary alignment. The API follows the skill's best practices for filters, loudness targets, and drift control.
💬 Integration Tip
Always confirm the TTS engine's native sample rate before resampling, and apply loudness normalization as the final step—re-normalize if you change tempo or duration afterward.
Scored Sep 29, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.