google-gemini-mediaUse the Gemini API (Nano Banana image generation, Veo video, Gemini TTS speech and audio understanding) to deliver end-to-end multimodal media workflows and code templates for "generation + understanding".
Install via ClawdBot CLI:
clawdbot install Xsir0/google-gemini-mediaGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
POST → https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:gCalls external URL not in known-safe list
https://generativelanguage.googleapis.com/v1beta/models/gemini-2.5-flash-image:gUses known external API (expected, informational)
googleapis.comAudited Apr 17, 2026 · audit v1.0
Generated Mar 1, 2026
Generate high-quality images of products for online stores, such as custom-designed items or food dishes, to enhance listings and marketing materials. This reduces the need for expensive photoshoots and allows rapid iteration on visual concepts.
Create narrated videos and images for e-learning platforms, explaining complex topics with AI-generated visuals and speech. This automates content production for courses, tutorials, and interactive learning modules.
Produce short promotional videos and images for social media campaigns, using text-to-video and image generation to quickly create engaging content. This streamlines ad creation and allows for personalized messaging at scale.
Analyze customer-uploaded images, videos, or audio to provide automated support, such as identifying product issues or transcribing support calls. This improves response times and reduces manual effort in helpdesk operations.
Edit existing images or extend videos for film, television, or digital media projects, using AI to enhance or modify visual content efficiently. This accelerates post-production workflows and reduces costs for studios.
Offer a cloud-based platform where users pay a monthly fee to access AI-powered media generation and understanding tools via API. This provides recurring revenue and scales with customer usage across various industries.
Charge customers based on the number of API calls or media processing tasks, such as per image generated or minute of video analyzed. This model appeals to businesses with variable needs and allows low entry costs.
License the skill package to other companies for integration into their own products, such as marketing tools or content management systems, with custom branding. This generates upfront licensing fees and ongoing support contracts.
💬 Integration Tip
Implement a file size threshold (e.g., 10MB) to automatically switch between inline and Files API modes for optimal performance and compliance with request limits.
Scored Apr 19, 2026
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
让 AI 代理根据对话内容自动选择最合适的模型。四层识别(系统过滤→关键词→指示词→语义相似度),四池架构(高速/智能/人文/代理),五分支路由,全自动 Fallback 回路。支持 trigger_groups_all 非连续词组命中。
OpenAI API integration — chat completions, embeddings, image generation, audio transcription, file management, fine-tuning, and assistants via the OpenAI RES...