Logo
ClawHub Skills Lib
HomeCategoriesUse CasesTrendingStatisticsBlog
HomeCategoriesUse CasesTrendingStatisticsBlog
ClawHub Skills Lib
ClawHub Skills Lib

Browse 50.000+ community-built AI agent skills for OpenClaw. Updated daily from clawhub.ai.

Explore

  • Home
  • Categories
  • Use Cases
  • Trending
  • Blog

Categories

  • Development
  • AI & Agents
  • Productivity
  • Communication
  • Data & Research
  • Business
  • Platforms
  • Lifestyle
  • Education
  • Design

Use Cases

  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • Knowledge Management
  • Email & Messaging
  • SEO & Content Marketing
  • Sales & CRM
  • Workflow Automation
  • Social Media
  • Chinese Platforms
  • E-Commerce
  • Education & Tutoring
  • HR & Recruiting
  • Legal & Compliance
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • See all use cases →
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • See all use cases →
© 2026 ClawHub Skills Lib. All rights reserved.Built with Next.js · Neon · Prisma
Home/Use Cases/🎙️ Voice & Audio AI/📝 Transcription & STT

📝 Transcription & STT AI Skills

Transcribe audio and video files to text with speaker labels and timestamps.

389 skillsPart of 🎙️ Voice & Audio AI
Lang:

389 skills found

Page 1 of 17

🎤Speech & Audio

AssemblyAI Transcriber

assemblyai-transcriber
xenofex7
v1.1.0
View Details

Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.

71
1.9k
4mo ago
🎤Speech & Audio

ElevenLabs Agents

elevenlabs-agents
pennyroyaltea
v1.0.0
View Details

Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.

10
4k
2
4mo ago
🎤Speech & Audio

yaps-transcription

yaps-transcription
yaps
v1.0.2
View Details

Turn an audio or video file into text with Yaps. Good for interviews, podcasts, and voice memos. New users: install Yaps and sign in.

+1
1
167
4d ago
🎤Speech & Audio

yaps-dictation

yaps-dictation
yaps
v1.0.2
View Details

Type with your voice using Yaps on desktop. Get set up, fix a problem, or recover a recent dictation. New users: install Yaps and sign in.

+1
1
161
4d ago
🎤Speech & Audio

Whisper Transcribe

whisper-transcribe
JosunLP
v1.0.0
View Details

Transcribe audio files to text using OpenAI Whisper. Supports speech-to-text with auto language detection, multiple output formats (txt, srt, vtt, json), batch processing, and model selection (tiny to large). Use when transcribing audio recordings, podcasts, voice messages, lectures, meetings, or any audio/video file to text. Handles mp3, wav, m4a, ogg, flac, webm, opus, aac formats.

0
2.5k
3
7mo ago
🎤Speech & Audio

Openai Tts.Bak 2026 01 28T18:01:23+10:30

openai-tts-bak-2026-01-28t18-01-23-10-30
nicoataiza
v1.0.0
View Details

Text-to-speech via OpenAI Audio Speech API.

0
2.2k
1
7mo ago
🎤Speech & Audio

Speech is Cheap Transcribe

asr
ilyakam
v1.2.0
View Details

Fast, affordable automatic speech-to-text transcription supporting 100 languages, speaker diarization, word timestamps, and customizable output formats.

0
3.9k
5
7mo ago
🎤Speech & Audio

🗣️ Edge-TTS Skill using uvx

edge-tts-uvx
al-one
v1.0.0
View Details

Text-to-speech conversion using `uvx edge-tts` for generating audio from text. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.

0
3.7k
3
7mo ago
🎤Speech & Audio

Elevenlabs AI

elevenlabs-ai
codedao12
v1.0.0
View Details

Access ElevenLabs APIs for text-to-speech, speech-to-speech, realtime speech-to-text, voice/model management, and dialogue workflows with direct HTTP calls.

0
2.3k
7mo ago
🎤Speech & Audio

Sound FX

sound-fx
javicasper
v0.1.1
View Details

Generate short sound effects via ElevenLabs SFX (text-to-sound). Use when you need SFX clips like applause, canned laughter, whooshes, ambience, or short stingers, and optionally convert to WhatsApp-friendly .ogg/opus.

0
3k
1
7mo ago
🎤Speech & Audio

Multimodal Base

yuyonghao-multimodal-base
yuyonghao-123
v0.1.0
View Details

Supports image understanding, OCR, speech-to-text, and text-to-speech synthesis with multi-voice and multimodal unified processing using OpenAI and Edge TTS.

0
682
4mo ago
🎤Speech & Audio

speech-translation

speech-translation
decin
v1.0.0
View Details

Build, adapt, or run an audio-processing workflow that takes spoken audio, transcribes it with Whisper or faster-whisper, translates the transcript using the...

0
799
4mo ago
🎤Speech & Audio

Voice Recognition

smart-voice-recognition
08jacky04
v1.1.0
View Details

Intelligent speech-to-text using local OpenAI Whisper (no API key needed, fully private). Use when you need to transcribe audio files, convert voice messages...

0
595
4mo ago
🎤Speech & Audio

Whisper GPU Audio Transcriber

whisper-gpu-transcriber-skill
allanmeng
v1.0.3
View Details

Convert audio to SRT subtitles using OpenAI Whisper with automatic GPU acceleration for Intel XPU / NVIDIA CUDA / AMD ROCm / Apple Metal. Ideal for content c...

+6
0
764
4mo ago
🎤Speech & Audio

Simple sound-to-text skill locally

sst-simple
lkisme
v1.0.2
View Details

Local speech-to-text using OpenAI Whisper. Use when the user needs to: (1) transcribe audio files to text, (2) convert voice messages to written content, (3)...

0
870
4mo ago
🎤Speech & Audio

llm-provider Whisper

openai-whisper-1-0-0
czubi1928
v1.0.0
View Details

基于Whisper CLI的本地语音转文字工具,无需API Key,支持中文交互和多场景自动化工作流应用。

0
69
7d ago
🎤Speech & Audio

SpeakNotes: YouTube, Audio & Document Summaries

speaknotes-youtube-audio-document-summarizer
JackLillie
v1.0.1
View Details

Use when OpenClaw needs to call SpeakNotes API routes directly using an API key and generate transcripts/summaries from YouTube URLs, media files, or documen...

+40
0
819
5mo ago
🎤Speech & Audio

Faster Whisper Local

faster-whisper-local
Damirikys
v1.0.0
View Details

Local speech-to-text using faster-whisper. High-performance transcription with GPU acceleration support. Includes word-level timestamps and distilled models....

0
1.7k
2
7mo ago
🎤Speech & Audio

Zhipu AI TTS

zhipu-tts
franklu0819-lang
v1.0.0
View Details

Text-to-speech conversion using Zhipu AI (BigModel) GLM-TTS model. Use when you need to convert text to audio files with various voice options. Supports Chin...

0
1.7k
1
7mo ago
🎤Speech & Audio

Voice-to-Protocol Transcriber

voice-to-protocol-transcriber-1
aipoch-ai
v1.0.0
View Details

Record experimental procedures and observations via voice commands during lab work. Real-time transcription for structured experiment documentation.

0
571
4mo ago
🎤Speech & Audio

Feishu Voice (ElevenLabs)

feishu-voice-elevenlabs
dongdongbear
v1.1.0
View Details

Send and receive voice messages on Feishu (Lark) using ElevenLabs TTS and STT. Activate when user asks to send a voice message on Feishu, or when receiving a...

0
1.2k
4mo ago
🎤Speech & Audio

Telegram Whisper Transcribe

telegram-whisper-transcribe
xela-io
v1.0.1
View Details

Standalone Telegram bot for voice message transcription via OpenAI Whisper API. No LLM overhead — audio goes directly to Whisper and text comes back in 2-5 s...

0
956
4mo ago
🎤Speech & Audio

audio-tools

emar-audio-tools
risehorizon
v1.0.0
View Details

音视频处理工具集。支持以下操作: - 从视频文件中提取音频并保存为 WAV 格式 - 对音频文件按指定开始时间和持续时长进行截取 - 播放指定的视频或音频文件(调用系统默认播放器) - 语音识别转文字(Whisper),输出 JSON 格式(含时间戳、置信度) - 提取音频/视频元数据(码率、采样率、时长、编码等...

0
655
1
4mo ago
🎤Speech & Audio

talkies

talkies
psyb0t
v1.3.17
View Details

Self-hosted OpenAI-compatible speech service. /v1/audio/transcriptions fronts 12 open ASR models (Whisper, Parakeet, Nemotron-3.5-ASR, Canary, Sherpa-ONNX, Vosk); /v1/audio/transcriptions/stream accepts live PCM over WebSocket. /v1/audio/speech fronts 3 TTS engines / 4 backends — Kokoro-82M (41 baked voices, PyTorch + ONNX runtimes), the CUDA-only Qwen3-TTS family (voice cloning, preset speakers, voice design), and the CUDA-only Chatterbox Turbo (English, 19 inline emotion tags, transcript-free cloning). Stereo diarization, URL fetching, six ASR/file-staging MCP tools, bearer auth.

0
1.8k
1mo ago
…

Other 🎙️ Voice & Audio AI Phases

🔊
Text-to-Speech
Convert text to natural-sounding speech with AI voice models and custom voice styles.
🌍
Audio Translation
Translate spoken content across languages — transcribe, translate, and re-synthesize.
🎚️
Audio Processing
Clean audio, remove noise, separate vocals, and process audio files at scale.