pdf-ocrPDF扫描件转Word文档。支持中文OCR识别,自动裁掉页眉页脚,保留插图,彩色章节封面页保留为图片。使用百度OCR API(免费额度1000次/月)。当用户要求把扫描PDF转成文字/Word时触发。
Install via ClawdBot CLI:
clawdbot install dadaniya99/pdf-ocrGrade Good — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated Mar 20, 2026
Researchers can convert scanned PDFs of historical or printed academic papers into editable Word documents for easier citation, analysis, and sharing. This is useful for digitizing archives or processing conference proceedings with mixed content like charts and images.
Law firms and legal departments can transform scanned legal documents, such as contracts or case files, into searchable and editable formats. This aids in document review, editing, and compliance checks, though manual verification may be needed for complex layouts.
Companies can digitize scanned business reports, financial statements, or meeting minutes into Word for data extraction, updates, and distribution. It streamlines workflows by making old documents accessible and editable, with compression for efficient sharing.
Publishers and libraries can convert scanned books or manuscripts into editable formats for reprinting, translation, or digital archiving. This handles mixed content like covers and illustrations, with compression to reduce file sizes for distribution.
Government agencies can digitize scanned public records, forms, or historical documents into Word for improved accessibility, editing, and long-term preservation. Manual checks may be required for complex pages to ensure accuracy.
Offer a free tier with 1000 pages per month using the Baidu OCR API, then charge for additional pages or premium features like faster processing or enhanced accuracy. This attracts users with basic needs and converts them to paying customers.
License the skill to businesses, such as legal firms or publishers, for bulk document processing with custom integrations, support, and higher usage limits. This targets organizations with regular digitization needs and offers scalable solutions.
Package the skill as a white-label tool for other software providers, like document management systems or educational platforms, to embed OCR functionality into their products. This generates revenue through partnership agreements or royalties.
💬 Integration Tip
Ensure proper API key management and test with sample PDFs to adjust cropping ratios for optimal OCR accuracy before full deployment.
Scored Apr 19, 2026
Create, inspect, and edit Microsoft Excel workbooks and XLSX files with reliable formulas, dates, types, formatting, recalculation, and template preservation...
基于智谱 GLM-OCR、GLM-4.7 及 GLM-4.6V 的多模态文档深度解析工具。 Use when: - 需要高精度提取文档(PDF/图片)中的表格并转换为 Markdown 格式 - 需要从文档页面中自动裁剪并提取插图、图表为独立文件 - 需要对提取的图表进行深度语义理解(基于 GLM-4.6V 视觉分析) - 需要对提取的表格数据进行逻辑分析(基于 GLM-4.7 文本分析) 核心架构: 1. 视觉提取:GLM-OCR 2. 语义理解:GLM-4.7 (纯文本/表格) + GLM-4.6V (多模态/图像)
Data analysis tool for Excel, CSV, Word, PDF, TXT, Markdown files. Use when user needs to analyze, summarize, or compare data from multiple files. Supports f...
PDF智能处理工具 v2.1 | PDF Smart Tool. 支持PDF转换、OCR识别、合并拆分、数字签名、批量处理、水印添加、加密解密。触发词:PDF、转换、识别。
Generate hand-drawn style diagrams, flowcharts, and architecture diagrams as PNG images from Excalidraw JSON
Convert public web pages into clean Markdown with markdown.new for AI workflows. Use when tasks require URL-to-Markdown conversion for summarization, RAG ing...