pocket-ttsGenerate high-quality English speech offline on CPU using 8 built-in voices or custom voice cloning with Kyutai's Pocket TTS model.
Install via ClawdBot CLI:
clawdbot install sherajdev/pocket-ttsGrade Good — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://huggingface.co/kyutai/pocket-ttsUses known external API (expected, informational)
arxiv.orgAudited Apr 16, 2026 · audit v1.0
Generated Mar 1, 2026
Enables text-to-speech in educational apps for students in low-connectivity areas, such as language learning platforms or e-readers. It supports voice cloning for personalized narration without internet dependency.
Integrates into on-premise customer service systems for businesses needing offline voice responses, like retail kiosks or factory assistance tools. Uses built-in voices or clones brand-specific tones.
Provides speech synthesis for accessibility features in software, such as screen readers for visually impaired users in offline environments. Runs on standard CPUs without GPU requirements.
Adds voice output to IoT devices like smart home assistants or industrial sensors that operate offline. Its low latency and CPU-only design suit resource-constrained hardware.
Supports local audio generation for content creators, such as podcasters or video editors needing voiceovers without cloud APIs. Voice cloning allows for custom character voices.
Offer a subscription-based service where businesses integrate Pocket TTS into their software for offline TTS capabilities. Revenue comes from monthly fees based on usage tiers or enterprise licenses.
Bundle the skill with hardware products like educational tablets or IoT devices that require offline speech synthesis. Revenue is generated through product sales and licensing agreements with manufacturers.
Provide a free basic version for developers, with premium features like advanced voice cloning or priority support. Revenue comes from paid upgrades and consulting services for custom integrations.
💬 Integration Tip
Ensure the Hugging Face model license is accepted before installation, and use the CLI for quick testing before Python API integration.
Scored Apr 19, 2026
Local speech-to-text with the Whisper CLI (no API key).
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
Turn your AI into JARVIS. Voice, wit, and personality — the complete package. Humor cranked to maximum.
Text-to-speech conversion using node-edge-tts npm package for generating audio from text. Supports multiple voices, languages, speed adjustment, pitch contro...
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...