ocr-local-v2Extract text from images using Tesseract.js OCR (100% local, no API key required). Supports Chinese (simplified/traditional) and English.
Install via ClawdBot CLI:
clawdbot install 15914355527/ocr-local-v2Grade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/naptha/tesseract.jsAudited Apr 16, 2026 · audit v1.0
Generated Oct 2, 2026
Finance teams can batch-process scanned invoices and receipts to extract vendor names, amounts, dates, and line items without manual data entry. The local OCR runs securely on-premises, keeping sensitive financial data private. Extracted text can be fed into accounting software or ERP systems for reconciliation.
Law firms can convert scanned contracts, court documents, and evidence images into searchable text in Chinese (simplified/traditional) and English. This enables keyword search, clause extraction, and e-discovery workflows without uploading confidential client materials to external cloud OCR services.
Construction and utility inspectors can photograph equipment labels, meters, and handwritten notes on-site, then OCR them into structured reports. The local processing works offline in remote areas without internet, and supports mixed Chinese-English labels common in international projects.
Sellers on cross-border platforms can extract text from supplier spec sheets, packaging photos, and product labels in Chinese and English. The OCR output can be automatically translated and reformatted for listings on Amazon, Shopify, or Taobao. This reduces manual copy-paste and speeds up catalog uploads.
A desktop or mobile helper app can let users snap a photo of any printed text—menus, signs, books—and hear it read aloud. Since OCR runs locally, it works without network and preserves user privacy. It supports bilingual environments, making it useful in Chinese-speaking regions.
Offer a containerized OCR microservice that companies can deploy behind their firewall. It provides a REST endpoint for internal apps to submit images and receive extracted text, with no per-request fees or cloud dependency. Support contracts and custom language model training can be upsold.
A cross-platform desktop app (Windows/macOS/Linux) that lets users drag-and-drop images for instant text extraction. The free tier handles basic images and limited pages per day; the paid tier unlocks batch processing, PDF multi-page support, and priority language downloads.
License the OCR engine as an embeddable JavaScript/npm package or SDK that other SaaS products can integrate into their document management or note-taking apps. Charge based on the number of active integrations or revenue share. Include private npm registry hosting and version updates.
💬 Integration Tip
Install tesseract.js via npm and call the script with node; for production use, pre-download language data and cache it locally to avoid first-run delays. Wrap the CLI in a simple HTTP server for easy integration with web or mobile apps.
Scored Oct 2, 2026
Fetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
Monitor RSS and Atom feeds for content research. Track blogs, news sites, newsletters, and any feed source. Use when monitoring competitors, tracking industr...
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Provides a personalized morning report with today's reminders, undone Notion tasks, and vault storage summary for daily planning.
Extract text from PDFs with OCR support. Perfect for digitizing documents, processing invoices, or analyzing content. Zero dependencies required.
Fetch scheduled economic events and data releases from the FMP API for specified dates, filtering by impact, country, and type, and output a chronological ma...