google-gemini-ttsGenerate spoken audio from text using Google's Gemini TTS models (default is Gemini 3.1 Flash TTS Preview, with fallback to Gemini 2.5 Flash/Pro preview TTS)...
Install via ClawdBot CLI:
clawdbot install shubhamsaboo/google-gemini-ttsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://ai.google.dev/gemini-api/docs/speech-generationUses known external API (expected, informational)
googleapis.comAudited Apr 17, 2026 · audit v1.0
Generated Aug 11, 2026
Contact centers use Gemini TTS to convert text-based support solutions into natural-sounding voice responses, reducing the need for human agents. The skill supports multiple languages and expressiveness, enabling a more human-like interaction in customer service bots.
Content creators can generate podcast-style two-speaker conversations from scripts, enabling rapid production of audio content without recording studios. The skill's multi-speaker mode and style control allow for dynamic and engaging podcasts.
Educational platforms use TTS to convert training materials and lessons into audio, providing accessible learning options. The ability to control style and voice supports different teaching personas and keeps learners engaged.
Apps and services for visually impaired users can leverage TTS to read web content, emails, or documents aloud. The multi-language support ensures accessibility across diverse user bases.
Smart home devices and virtual assistants use Gemini TTS to provide natural, context-aware voice replies. Style control enables the assistant to convey emotions, making interactions feel more human and personal.
Offer text-to-speech as a service with tiered pricing based on usage (number of characters or minutes of audio). Businesses pay monthly to integrate TTS into their products.
Expose the TTS functionality via a REST API and charge per API call or per second of audio generated. Developers integrate the API into their apps and services.
Build a platform where users create audio content (podcasts, audio articles) using TTS and monetize through premium features (advanced voices, longer audio, commercial usage rights).
💬 Integration Tip
Use the provided CLI script to get started quickly; export your GEMINI_API_KEY and ensure curl, jq, base64, and ffmpeg are installed. For programmatic integration, wrap the script in your application's backend to handle text input and output WAV file paths.
Scored Aug 11, 2026
Transcribe audio files with speaker diarization (who speaks when). Supports 100+ languages, automatic language detection, and timestamps. Use for meetings, interviews, podcasts, or voice messages. Requires AssemblyAI API key.
使用 poocr 库识别发票并导出 Excel。当用户需要识别增值税发票、批量处理发票文件或提取发票信息到 Excel 时调用此技能。
Use when the user has audio or video and wants a timestamped transcript (SRT) in the source language. Routes by source language — Chinese defaults to Volcano...
Create, manage, and deploy ElevenLabs conversational AI agents. Use when the user wants to work with voice agents, list their agents, create new ones, or manage agent configurations.
Voice note transcription and archival for OpenClaw agents. Powered by Deepgram Nova-3. Transcribes audio messages, saves both audio files and text transcript...
Text-to-Speech (TTS) and Speech-to-Text (ASR) using coze-coding-dev-sdk. Returns results directly to stdout.