Logo
ClawHub Skills Lib
HomeCategoriesUse CasesTrendingStatisticsBlog
HomeCategoriesUse CasesTrendingStatisticsBlog
ClawHub Skills Lib
ClawHub Skills Lib

Browse 50.000+ community-built AI agent skills for OpenClaw. Updated daily from clawhub.ai.

Explore

  • Home
  • Categories
  • Use Cases
  • Trending
  • Blog

Categories

  • Development
  • AI & Agents
  • Productivity
  • Communication
  • Data & Research
  • Business
  • Platforms
  • Lifestyle
  • Education
  • Design

Use Cases

  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • Knowledge Management
  • Email & Messaging
  • SEO & Content Marketing
  • Sales & CRM
  • Workflow Automation
  • Social Media
  • Chinese Platforms
  • E-Commerce
  • Education & Tutoring
  • HR & Recruiting
  • Legal & Compliance
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • Crypto & Web3
  • Real-Time Web Search
  • News & Media Monitoring
  • Academic Research
  • Data & Analytics
  • AI Image Generation
  • Voice & Audio AI
  • AI Video Creation
  • Content Writing
  • Task & Project Management
  • See all use cases →
  • AI Code Generation
  • Code Review & Testing
  • DevOps & Cloud
  • Security & Compliance
  • Build an AI Agent
  • Agent Memory & RAG
  • Multi-Agent Orchestration
  • Browser & Web Automation
  • Financial & Market Data
  • See all use cases →
© 2026 ClawHub Skills Lib. All rights reserved.Built with Next.js · Neon · Prisma
Home/Use Cases/🎙️ Voice & Audio AI/🎚️ Audio Processing

🎚️ Audio Processing AI Skills

Clean audio, remove noise, separate vocals, and process audio files at scale.

64 skillsPart of 🎙️ Voice & Audio AI
Lang:

64 skills found

Page 1 of 3

🎤Speech & Audio

yaps-audio-cleaner

yaps-audio-cleaner
yaps
v1.0.2
View Details

Reduce noise, hiss, and static in a speech recording with Yaps. Save a separate clean file. New users: install Yaps and sign in.

+1
1
169
4d ago
🎤Speech & Audio

Whisper Transcribe

whisper-transcribe
JosunLP
v1.0.0
View Details

Transcribe audio files to text using OpenAI Whisper. Supports speech-to-text with auto language detection, multiple output formats (txt, srt, vtt, json), batch processing, and model selection (tiny to large). Use when transcribing audio recordings, podcasts, voice messages, lectures, meetings, or any audio/video file to text. Handles mp3, wav, m4a, ogg, flac, webm, opus, aac formats.

0
2.5k
3
7mo ago
🎤Speech & Audio

ElevenLabs

elevenlabs-api
byungkyu
v1.2.5
View Details

ElevenLabs API integration with managed authentication. AI-powered text-to-speech, voice cloning, sound effects, and audio processing. Use this skill when users want to generate speech from text, clone voices, create sound effects, or process audio. For other third party apps, use the api-gateway skill (https://clawhub.ai/byungkyu/api-gateway). Calls run through the `maton` CLI with OAuth login, or over raw HTTP with a Maton API key where the CLI cannot be installed. Every call is authenticated as the user's connection and reaches only what that connection's authorization allows, which the provider enforces on every request; the endpoints documented here are the ones this skill uses, and any other endpoint of this app needs the user to ask for it by name. Default to read and list calls, and confirm every write or new connection with the user. This file also documents the three constructs that turn an ElevenLabs connection into automation, in the order they are used: the connection (the first step), a hosted function that runs an ElevenLabs action through the Maton SDK, and a trigger that calls that function on a schedule or on an event. Those sections are the platform's own reference text, shared with the api-gateway skill, with ElevenLabs examples; they add no ElevenLabs capability - ElevenLabs is not an event source, a trigger cannot read ElevenLabs data, and the files under `references/<source>/triggers.md` are the platform's event catalogues for the sources Maton offers (time, Calendly, GitHub, Gmail, HubSpot, Linear, Notion, Slack, Stripe).

0
3.7k
3
5d ago
🎤Speech & Audio

Multimodal Base

yuyonghao-multimodal-base
yuyonghao-123
v0.1.0
View Details

Supports image understanding, OCR, speech-to-text, and text-to-speech synthesis with multi-voice and multimodal unified processing using OpenAI and Edge TTS.

0
682
4mo ago
🎤Speech & Audio

speech-translation

speech-translation
decin
v1.0.0
View Details

Build, adapt, or run an audio-processing workflow that takes spoken audio, transcribes it with Whisper or faster-whisper, translates the transcript using the...

0
799
4mo ago
🎤Speech & Audio

Text to Audio GPT

text-to-audio-gpt
ciklopentan
v1.0.2
View Details

GPT-only: озвучивает текст и документы выразительным голосом и собирает проверенный MP3; подходит для чтения вслух, аудиокниг и озвучки. Требует GPT browser control и FFmpeg/FFprobe, не нативен для OpenClaw; Android ChatGPT требует отдельного импорта ZIP.

0
117
5d ago
🎤Speech & Audio

SpeakNotes: YouTube, Audio & Document Summaries

speaknotes-youtube-audio-document-summarizer
JackLillie
v1.0.1
View Details

Use when OpenClaw needs to call SpeakNotes API routes directly using an API key and generate transcripts/summaries from YouTube URLs, media files, or documen...

+40
0
819
5mo ago
🎤Speech & Audio

pixelhub-api-tools

pixelhub-api-tools
leik1000
v1.0.7
View Details

Use for Pixelhub API direct calls when users need image generation/editing, video generation/post-processing, or audio/music generation.

0
972
4mo ago
🎤Speech & Audio

AudioPod

audiopod
Rakesh1002
v1.2.3
View Details

Use AudioPod AI's API for audio processing tasks including AI music generation (text-to-music, text-to-rap, instrumentals, samples, vocals), stem separation, text-to-speech, noise reduction, speech-to-text transcription, speaker separation, and media extraction. Use when the user needs to generate music/songs/rap from text, split a song into stems/vocals/instruments, generate speech from text, clean up noisy audio, transcribe audio/video, or extract audio from YouTube/URLs. Requires AUDIOPOD_API_KEY env var or pass api_key directly.

0
4.3k
3
7mo ago
🎤Speech & Audio

Voice Note To Midi

voice-note-to-midi
danbennettuk
v0.1.0
View Details

Convert voice notes, humming, and melodic audio recordings to quantized MIDI files using ML-based pitch detection and intelligent post-processing

0
2.7k
4mo ago
🎤Speech & Audio

Whisper Transcriber

whisper-transcriber
vvusu
v1.0.0
View Details

Offline speech-to-text (ASR) using whisper.cpp (whisper-cli) + ffmpeg. Supports batch transcription, timestamps, SRT/TXT/JSON outputs, and model download. Cr...

0
1k
1
5mo ago
🎤Speech & Audio

AssemblyAI advanced speech transcription

assemblyai-transcribe
tristanmanchester
v1.0.1
View Details

Transcribe, diarise, translate, post-process, and structure audio/video with AssemblyAI. Use this skill when the user wants AssemblyAI specifically, needs hi...

0
4.1k
3
4mo ago
🎤Speech & Audio

Parakeet Stt

parakeet-stt
carlulsoe
v1.1.0
View Details

Local speech-to-text with NVIDIA Parakeet TDT 0.6B v3 (ONNX on CPU). 30x faster than Whisper, 25 languages, auto-detection, OpenAI-compatible API. Use when transcribing audio files, converting speech to text, or processing voice recordings locally without cloud APIs.

0
3.2k
1
4mo ago
🎤Speech & Audio

Jarvis Vocal

jarvis-vocal
kishen35
v1.0.0
View Details

Authentic J.A.R.V.I.S. voice synthesis using Piper TTS with HuggingFace-trained model. Generates movie-accurate voice locally and can push to connected Andro...

+1
0
784
4mo ago
🎤Speech & Audio

Text to speech using the default macOS "say" command. No need for 3rd party APIs or models. Supports many languages. Also, Trinoids!

macos-say
zviratko
v0.0.2
View Details

Local text-to-speech using macOS `say` + ffmpeg for Telegram/Matrix voice messages

0
570
4mo ago
🎤Speech & Audio

mmVoiceMaker

mm-voice-maker
blue-coconut
v1.0.1
View Details

Enables voice synthesis, voice cloning, voice design, and audio post-processing using MiniMax Voice API and FFmpeg. Use when converting text to speech, creat...

0
1.5k
3
5mo ago
🎤Speech & Audio

podcast-transcribe

podcast-transcribe
dairui1
v1.0.0
View Details

Download podcast audio from RSS feeds and transcribe to text using AuralWise API. This skill should be used when the user wants to download podcast episodes, convert podcast audio to text transcripts, or batch-process a podcast library for searchable content. Triggers include downloading podcasts, podcast transcription, audio-to-text conversion, RSS feed downloading, or any request involving podcast audio acquisition and speech-to-text conversion. Covers RSS feed discovery, audio downloading, AuralWise API transcription, and AI-generated content overviews with book references, key concepts, and searchable keywords.

0
16
2mo ago
🎤Speech & Audio

音乐生成 Suno Music

dlazy-suno-music
dlazyai
v1.3.18
View Details

Suno music generation model. Supports inspiration mode (auto lyrics) and custom mode (manual lyrics), generating music with or without vocals. Suno 音乐生成模型。支持灵感模式(自动作词)和自定义模式(手动填词),可生成包含人声或纯器乐的音乐。

0
4.1k
7d ago
🎤Speech & Audio

purevocals-uvr-automator

purevocals-uvr-automator
wangminrui2022
v1.0.5
View Details

当用户想要**一键批量从音频文件中提取超干净纯人声(干声 / Vocals Only)**、去除伴奏/背景音乐时,自动调用此技能。 一键音频人声分离工具。专门从音频文件(.mp3/.wav/.flac等)中提取超干净干声(Acapella)或去除背景音制作伴奏。 核心用途:支持单个音频文件或整个文件夹批量处理(.mp3/.wav/.flac 等格式),输出高质量无杂音干声,自动在输入同级创建 [输入文件夹]_vocals 文件夹,完美保留原目录结构。 高频触发场景包括: - 翻唱练习、翻唱视频制作、B站/抖音/小红书演唱素材清洗 - 卡拉OK 伴奏制作(只保留人声) - 音乐制作中的人声分

0
988
2mo ago
🎤Speech & Audio

audiobook

pg-essay-to-audiobook-audiobook
lnj22
v0.1.0
View Details

Create audiobooks from web content or text files. Handles content fetching, text processing, and TTS conversion with automatic fallback between ElevenLabs, O...

0
546
4mo ago
🎤Speech & Audio

Kie Audio Generator

kie-audio
benhuebner01
v1.0.0
View Details

Generate music and audio via Kie.ai's Suno gateway (V3.5 through V5.5). Use for background tracks, instrumental beds, full songs with vocals, or extending ex...

0
602
4mo ago
🎤Speech & Audio

UpInvoice: Document AI

upinvoice
upinvoice
v1.0.0
View Details

Extract structured JSON data from invoice PDFs or images using UpInvoice.eu AI for fast, accurate automation of invoice processing in ERP systems.

0
566
4mo ago
🎤Speech & Audio

Record2Note

record2note
soullhcn
v1.1.0
View Details

Use when converting voice recordings, voice memos, or speech transcripts into structured Markdown notes, setting up recording-folder monitoring, or processing pending transcription files.

+1
0
448
3mo ago
🎤Speech & Audio

voice

godfery-voice
godferylindsay
v1.0.0
View Details

Enables real-time voice chat in Discord channels with speech-to-text, Claude AI processing, and text-to-speech playback supporting barge-in and auto-reconnect.

0
552
4mo ago

Other 🎙️ Voice & Audio AI Phases

🔊
Text-to-Speech
Convert text to natural-sounding speech with AI voice models and custom voice styles.
📝
Transcription & STT
Transcribe audio and video files to text with speaker labels and timestamps.
🌍
Audio Translation
Translate spoken content across languages — transcribe, translate, and re-synthesize.