multimodal使用GLM-4.6V模型进行多模态内容理解(图片、视频、文档)
Install via ClawdBot CLI:
clawdbot install tridefender/multimodalGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://open.bigmodel.cn/api/paas/v4/chat/completionsAudited Apr 17, 2026 · audit v1.0
Generated Mar 20, 2026
Automatically extract product attributes, detect objects, and perform OCR on product images to generate detailed descriptions and tags. This helps in catalog management and enhances search functionality for online retailers.
Analyze educational videos and PDF documents to create summaries, identify key concepts, and extract complex tables for study aids. Useful for online learning platforms to improve content accessibility and engagement.
Process images and videos from social media or news sources to detect scenes, objects, and text for content moderation and regulatory compliance. Assists media companies in monitoring brand safety and adherence to guidelines.
Parse medical PDFs and forms, extracting data from complex tables and performing OCR on scanned documents to digitize patient records. Supports healthcare providers in automating administrative tasks and improving data accuracy.
Analyze surveillance video feeds to generate summaries, detect key events, and identify objects or text in images for threat assessment. Used by security firms to enhance monitoring efficiency and incident response.
Offer the skill as a cloud-based API, charging per request or through subscription tiers based on usage volume. This model targets developers and businesses needing scalable multimodal analysis without infrastructure management.
Provide customized on-premise or private cloud deployments with dedicated support and integration services for large organizations. This ensures data privacy and tailored solutions for industries like healthcare or finance.
Embed the skill into existing platforms such as content management systems or analytics tools, generating revenue through partnership agreements or revenue-sharing models. This expands reach by leveraging established user bases.
💬 Integration Tip
Ensure the ZHIPU_API_KEY is securely set as an environment variable and validate input URLs for accessibility before processing to avoid API errors.
Scored Apr 19, 2026
Use CodexBar CLI local cost usage to summarize per-model usage for Codex or Claude, including the current (most recent) model or a full model breakdown. Trigger when asked for model-level usage/cost data from codexbar, or when you need a scriptable per-model summary from codexbar cost JSON.
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
Sync OpenRouter models used by OpenClaw into this installation's config. Fetches the OpenClaw app leaderboard from OpenRouter, verifies model IDs against the...