Logo
ClawHub Skills Lib
HomeCategoriesUse CasesTrendingStatisticsBlog
HomeCategoriesUse CasesTrendingStatisticsBlog
ClawHub Skills Lib
ClawHub Skills Lib

Browse 50.000+ community-built AI agent skills for OpenClaw. Updated daily from clawhub.ai.

Explore

  • Home
  • Categories
  • Use Cases
  • Trending
  • Blog

Categories

  • Development
  • AI & Agents
  • Productivity
  • Communication
  • Data & Research
  • Business
  • Platforms
  • Lifestyle
  • Education
  • Design

Use Cases

  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • Knowledge Management
  • Email & Messaging
  • SEO & Content Marketing
  • Sales & CRM
  • Workflow Automation
  • Social Media
  • Chinese Platforms
  • E-Commerce
  • Education & Tutoring
  • HR & Recruiting
  • Legal & Compliance
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • See all use cases →
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • See all use cases →
© 2026 ClawHub Skills Lib. All rights reserved.Built with Next.js · Neon · Prisma
Home/🤖 AI & Agents/🎤 Speech & Audio

🎤 Speech & Audio AI Skills

1,304 AI agent skills for Speech & Audio. Part of the 🤖 AI & Agents category.

Speech & Audio Skills

Lang:

1,304 skills found

Page 1 of 55

🎤Speech & Audio

Openai Whisper

openai-whisper
steipete
Av1.0.0
View Details

Local speech-to-text with the Whisper CLI (no API key).

3k
85.2k
328
6mo ago
🎤Speech & Audio

Sag

sag
steipete
Av1.0.0
View Details

ElevenLabs text-to-speech with mac-style say UX.

0
28.7k
29
6mo ago
🎤Speech & Audio

Openai Whisper Api

openai-whisper-api
steipete
Av1.0.0
View Details

Transcribe audio via OpenAI Audio Transcriptions API (Whisper).

0
27.6k
52
3mo ago
🎤Speech & Audio

Edge TTS

edge-tts
i3130002
Av2.0.0
View Details

Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch control, and subtitle generation. Use when: (1) User requests audio/voice output with the "tts" trigger or keyword. (2) Content needs to be spoken rather than read (multitasking, accessibility, driving, cooking). (3) User wants a specific voice, speed, pitch, or format for TTS output.

0
22.4k
33
6mo ago
🎤Speech & Audio

ClawCall

clawcall-dev
clawcall-dev
Av1.0.6
View Details

Use when the user wants an AI agent to place a US phone call, call a business, handle hold or phone menus, confirm/reschedule/cancel/book/follow up/check an...

0
19.3k
10
3mo ago
🎤Speech & Audio

cellcog

cellcog
nitishgargiitd
Av2.0.21
View Details

Any-to-any AI sub-agent — research, images, video, audio, music, podcasts, avatars, voice cloning, documents, spreadsheets, dashboards, 3D models, diagrams, and code in one request. Agent-to-agent protocol with multi-step iteration for high accuracy. #1 on DeepResearch Bench (Apr 2026) — deep reasoning meets all modalities, so all your work gets done, not just code.

0
17.8k
8
15d ago
🎤Speech & Audio

Local Whisper

local-whisper
araa47
Av1.0.0
View Details

Local speech-to-text using OpenAI Whisper. Runs fully offline after model download. High quality transcription with multiple model sizes.

0
13.5k
12
6mo ago
🎤Speech & Audio

Voice Wake Say

voice-wake-say
xadenryan
Av1.0.1
View Details

Speak responses aloud on macOS using the built-in `say` command when user input indicates Voice Wake/voice recognition (for example, messages starting with "User talked via voice recognition on <device>").

0
10.9k
5
3mo ago
🎤Speech & Audio

Faster Whisper

faster-whisper
theplasmak
Av1.5.1
View Details

Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables ~20x realtime transcription. SRT...

0
8.4k
5
3mo ago
🎤Speech & Audio

OpenAI TTS

openai-tts
pors
Bv1.0.0
View Details

Text-to-speech via OpenAI Audio Speech API.

0
7.9k
6
6mo ago
🎤Speech & Audio

Kokoro TTS

kokoro-tts
edkief
Av0.1.0
View Details

Generate spoken audio from text using the local Kokoro TTS engine. Use when the user asks to "say" something, requests a voice message, or wants text converted to speech.

0
7.9k
2
6mo ago
🎤Speech & Audio

ElevenLabs Voices

elevenlabs-voices
robbyczgw-cla
Sv2.1.6
View Details

High-quality voice synthesis with 18 personas, 32 languages, sound effects, batch processing, and voice design using ElevenLabs API.

0
7.8k
16
3mo ago
🎤Speech & Audio

Mac TTS

mac-tts
kalijason
Bv1.0.0
View Details

Text-to-speech using macOS built-in `say` command. Use for voice notifications, audio alerts, reading text aloud, or announcing messages through Mac speakers. Supports multiple languages including Chinese (Mandarin), English, Japanese, etc.

0
7.6k
2
6mo ago
🎤Speech & Audio

Audio Generation

audio-generation-cellcog
nitishgargiitd
Av1.0.17
View Details

AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Three voice providers (OpenAI, ElevenLabs, MiniMax). Professional audio production from text prompts.

0
7.3k
4
15d ago
🎤Speech & Audio

Voice Transcribe

voice-transcribe
darinkishore
Av1.0.1
View Details

Transcribe audio files using OpenAI's gpt-4o-mini-transcribe model with vocabulary hints and text replacements. Requires uv (https://docs.astral.sh/uv/).

0
6.8k
13
4mo ago
🎤Speech & Audio

Audio Cog

audio-cog
nitishgargiitd
Av1.0.12
View Details

AI audio generation and text-to-speech powered by CellCog. Voiceover, narration, voice cloning, avatar voices, sound effects, music, podcasts, dialogue. Thre...

0
6.3k
4
4mo ago
🎤Speech & Audio

Audio

audio
ivangdavila
Bv1.0.1
View Details

Process, enhance, and convert audio files with noise removal, normalization, format conversion, transcription, and podcast workflows.

0
5.9k
2
3mo ago
🎤Speech & Audio

speech-recognition

speech-recognition
demo112
Bv1.0.1
View Details

通用语音识别 Skill。支持多种音频格式(ogg/mp3/wav/m4a),使用硅基流动 SenseVoice API 进行语音转文字。当用户发送语音消息、音频文件,或需要转录音频时触发。

0
5.9k
3
4mo ago
🎤Speech & Audio

macOS Local Voice

macos-local-voice
strrl
Bv1.0.0
View Details

Local STT and TTS on macOS using native Apple capabilities. Speech-to-text via yap (Apple Speech.framework), text-to-speech via say + ffmpeg. Fully offline, no API keys required. Includes voice quality detection and smart voice selection.

0
5.8k
1
3mo ago
🎤Speech & Audio

Elevenlabs

elevenlabs
odrobnik
Av1.3.4
View Details

Text-to-speech, sound effects, music generation, voice management, and quota checks via the ElevenLabs API. Use when generating audio with ElevenLabs or mana...

0
5.8k
2
6mo ago
🎤Speech & Audio

Music Generation

music-generation-cellcog
nitishgargiitd
Bv1.0.15
View Details

AI music generation powered by CellCog. Original instrumental and vocal tracks, 5 seconds to 10 minutes. Cinematic scores, background tracks, podcast intros, game soundtracks, ambient soundscapes, jingles, lo-fi beats, orchestral compositions, songs with lyrics. Royalty-free.

0
5.5k
4
15d ago
🎤Speech & Audio

Jarvis Voice

jarvis-voice
globalcaos
Av2.2.1
View Details

Turn your AI into JARVIS. Voice, wit, and personality — the complete package. Humor cranked to maximum.

0
5.4k
4
5mo ago
🎤Speech & Audio

Alexa CLI

alexa-cli
buddyh
Av1.3.0
View Details

Control Amazon Alexa devices and smart home via the `alexacli` CLI. Use when a user asks to speak/announce on Echo devices, control lights/thermostats/locks, send voice commands, or query Alexa.

0
5.4k
14
4mo ago
🎤Speech & Audio

Tts

tts
AMSTKO
Bv1.0.0
View Details

Convert text to speech using Hume AI (or OpenAI) API. Use when the user asks for an audio message, a voice reply, or to hear something "of vive voix".

0
5.3k
1
6mo ago
…

More in 🤖 AI & Agents

🛡️
Agent Security
0 skills
🧠
LLMs & Model APIs
1127 skills
🤖
Agent Frameworks
7944 skills
🧠
Agent Memory
449 skills
🔄
Agent Self-Improvement
319 skills
⚙️
AI Tools & Utilities
1329 skills
🖼️
Image Generation
1966 skills
🎬
Video Generation
548 skills
⚡
Automation & Workflows
318 skills
💬
Chatbots & Assistants
1127 skills
📝
Prompt & Config
413 skills