vision-helperAnalyze images using local or cloud vision models via Ollama to identify content, UI elements, screenshots, or extract text with OCR support.
Install via ClawdBot CLI:
clawdbot install ravenquasar/vision-helperGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
exec(Calls external URL not in known-safe list
http://localhost:11434/api/chat`Audited Apr 30, 2026 · audit v1.0
Generated Oct 6, 2026
QA engineers capture browser or desktop screenshots of application bugs and pipe them into vision-helper for automatic description, UI element identification, and OCR of error messages. The structured analysis is attached directly to bug tickets, reducing manual documentation time.
Developers building game-playing agents screenshot the current frame, call vision-helper to identify board state, UI buttons, or resource counters, then execute clicks or keyboard input based on the analysis. This creates a perception-action loop for card games, chess, and mobile puzzle apps.
Teams digitize scanned invoices, receipts, ID cards, or forms by feeding images through vision-helper with a custom prompt to extract text and fields. Cloud vision models handle messy scans and multiple languages including Chinese, outputting structured text for downstream processing.
Content managers and accessibility teams auto-generate descriptive alt text for images uploaded to websites, CMSs, or e-commerce catalogs. The skill analyzes each image and returns a natural-language description that is inserted into HTML alt attributes at scale.
IT operations schedule periodic desktop screenshots and run vision-helper to detect visual changes, modal dialogs, or unexpected error windows on kiosks and lab machines. Alerts fire when the AI describes an anomaly, enabling unattended monitoring.
The skill script itself is free and works with local Ollama models, but advanced cloud vision models like kimi-k2.6 or qwen3.5 provide higher accuracy and are billed per API call through an Ollama-hosted or vendor endpoint. Users adopt free for prototyping and convert to paid cloud usage as volume grows.
A hosted service wraps the vision-helper workflow with a REST endpoint, authentication, and queues, so customers avoid running Ollama locally. Pricing tiers are based on monthly image volumes with SLAs for latency and uptime.
Packaged solutions for niches such as game bots, invoice OCR, or UI regression testing bundle the skill with custom prompts, dashboards, and integrations. Customers pay for the outcome-oriented workflow rather than raw image analysis.
💬 Integration Tip
Set the exec timeout to 120–180 seconds whenever calling the script, since cloud vision models routinely take 40–120 seconds and will fail under shorter limits. Use the VISION_MODEL environment variable to switch between a fast local model for privacy-sensitive work and a cloud model for higher-accuracy tasks.
Scored Oct 6, 2026
Fetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
Monitor RSS and Atom feeds for content research. Track blogs, news sites, newsletters, and any feed source. Use when monitoring competitors, tracking industr...
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Provides a personalized morning report with today's reminders, undone Notion tasks, and vault storage summary for daily planning.
Extract text from PDFs with OCR support. Perfect for digitizing documents, processing invoices, or analyzing content. Zero dependencies required.
Fetch scheduled economic events and data releases from the FMP API for specified dates, filtering by impact, country, and type, and output a chronological ma...