Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Convert text to natural speech, transcribe audio to text, and translate spoken content with AI.
These skills handle the full audio pipeline — synthesizing voices with ElevenLabs and OpenAI TTS, transcribing recordings with Whisper, translating spoken content across languages, and processing audio files. Used by podcasters, developers, and accessibility teams.
Convert text to natural-sounding speech with AI voice models and custom voice styles.
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Turn an audio or video file into text with Yaps. Good for interviews, podcasts, and voice memos. New users: install Yaps and sign in.
Turn text into spoken audio with Yaps. Choose a supported voice and language, then save the file. New users: install Yaps and sign in.
Type with your voice using Yaps on desktop. Get set up, fix a problem, or recover a recent dictation. New users: install Yaps and sign in.
Quick install — most popular voice & audio ai skill:
clawdbot install pennyroyaltea/elevenlabs-agents563 skills found
Page 1 of 24
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
演讲稿(TED风格)、婚礼致辞、商务演讲、励志演讲、祝酒词、演讲大纲。Speech writing for TED-style talks, wedding speeches, business presentations, motivational speeches, toasts, and outlines....
Text-to-speech, speech-to-text, voice conversion, and audio processing using EachLabs AI models. Supports ElevenLabs TTS, Whisper transcription with diarization, and RVC voice conversion. Use when the user needs TTS, transcription, or voice conversion.
语音笔记转文字工具 v2.1 | Voice Note Transcriber. 支持多语言识别、实时转写、说话人识别、智能摘要、音频降噪、离线识别。触发词:转写、识别、语音。
Transcribe audio via Deepgram Nova-3 API (5.26% WER, 40x faster than Whisper, built-in speaker diarization). Use when user asks to transcribe audio, podcasts...
Check a YouTube channel's newest uploads for free via the BulkTranscripts API and fetch transcripts only for videos that are new. Use when the user asks what a channel posted recently, wants alerts or digests for new uploads, or wants to track competitors' channels. Requires a free BulkTranscripts API key in BULKTRANSCRIPTS_API_KEY, created at https://bulktranscripts.co/app?tab=mcp (Google sign-in, 30 free credits, no card).
Get the transcript of a single YouTube video (or Short, or TikTok video) as clean text, timestamped segments, or SRT/VTT/Markdown via the BulkTranscripts API. Use when the user shares a video link and wants it summarized, quoted, translated, fact-checked or turned into notes. Requires a free BulkTranscripts API key in BULKTRANSCRIPTS_API_KEY, created at https://bulktranscripts.co/app?tab=mcp (Google sign-in, 30 free credits, no card).
Build and manage Voice AI agents using Vapi, Bland.ai, or Retell. Create agents, configure voices, set prompts, make outbound calls, and retrieve transcripts...
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Make a meeting transcript with speaker labels in Yaps. Correct it, export it, or ask questions. New users: install Yaps and sign in.
Save the sound from a video as MP3, WAV, or M4A with Yaps. New users: install Yaps and sign in.
Translate text or subtitle files with Yaps. Keep the original and preserve subtitle timing. New users: install Yaps and sign in.
Turn an audio or video file into text with Yaps. Good for interviews, podcasts, and voice memos. New users: install Yaps and sign in.
Turn text into spoken audio with Yaps. Choose a supported voice and language, then save the file. New users: install Yaps and sign in.
Make a timed SRT subtitle file from audio or video with Yaps. New users: install Yaps and sign in.
Type with your voice using Yaps on desktop. Get set up, fix a problem, or recover a recent dictation. New users: install Yaps and sign in.
Add captions to a video with Yaps. Review the words and timing, choose a style, then export. New users: install Yaps and sign in.
Reduce noise, hiss, and static in a speech recording with Yaps. Save a separate clean file. New users: install Yaps and sign in.
Use Yaps for voice, audio, video, images, translation, and memory when no focused skill fits. New users: install Yaps and sign in.
语音笔记转文字工具 Pro | 支持多语言语音识别、实时转写、会议纪要生成。
将已完成的点播视频转写为文字与结构化文案。当用户提供本地视频文件或抖音、小红书等公网视频链接,要求视频转文字、视频转稿、字幕提取、视频总结、会议纪要、课程拆解、二创改写时使用。不支持实时直播流、需登录的加密视频与纯音乐无人声视频;转写后可按自定义 Prompt 生成总结、金句、分镜头、翻译等
读取抖音视频的内容——元信息、官方 AI 章节要点,以及通过逐帧截图 + 字幕 OCR 得到完整口播讲稿。用户分享抖音链接(v.douyin.com / douyin.com/video/xxx)并希望了解视频讲了什么、提取文案、拿到文字稿时使用。触发词:抖音链接、抖音视频、这个视频讲了什么、提取视频文案、抖音视频转文字、视频字幕、看看这个视频、视频内容。
Generate music via Suno with the local browser-backed flow. Use when the user wants Suno songs, instrumental tracks, lyric-based songs, Suno credit checks, o...
Skills integrate with ElevenLabs, OpenAI TTS, Murf, Play.ht, Azure Cognitive Services, and Google Cloud TTS — covering natural voices, voice cloning, and custom voice styles.
Yes. Whisper-based skills handle long-form audio by chunking and processing in segments, with speaker diarization and timestamp output. Cloud-based skills handle files up to several hours.