pdf-ocr-skill支持双引擎的PDF OCR识别技能,可从影印版PDF文件和图片文件中提取文字内容
Install via ClawdBot CLI:
clawdbot install yejinlei/pdf-ocr-skillGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Sends data to undocumented external endpoint (potential exfiltration)
Report → https://github.com/yejinlei/pdf-ocr-skill/issuesCalls external URL not in known-safe list
https://github.com/yejinlei/pdf-ocr-skillAudited Apr 17, 2026 · audit v1.0
Generated Mar 22, 2026
Law firms and legal departments use this skill to extract text from scanned contracts, agreements, and court documents for digital archiving and analysis. It supports both local and cloud OCR engines to handle varying quality scans while maintaining text structure.
Researchers and librarians process scanned PDFs of historical books, journals, and manuscripts to create searchable digital copies. The dual-engine capability allows for cost-effective local processing with high-accuracy cloud options for complex texts.
Healthcare providers digitize patient records, prescriptions, and lab reports from scanned images or PDFs for electronic health record systems. The skill handles various image formats and ensures accurate text extraction for compliance and data analysis.
Companies automate the extraction of text from scanned invoices, receipts, and financial documents for accounting and expense tracking. The skill's batch processing feature enables efficient handling of large volumes of documents.
Publishers and media agencies convert scanned magazines, newspapers, and brochures into editable digital formats for repurposing or online distribution. Support for Chinese and English text makes it suitable for multilingual publications.
Offer free usage with the local RapidOCR engine for basic needs, and charge subscription fees for access to the high-accuracy Silicon Flow API engine. This model attracts users with no-cost entry and monetizes advanced features.
Sell custom licenses to large organizations for on-premise deployment with enhanced support, security features, and integration capabilities. This model targets industries like legal and healthcare with strict data privacy requirements.
Provide the Silicon Flow OCR engine as a standalone API service for developers to integrate into their applications, charging based on usage volume such as per-page or per-API call. This model scales with customer demand.
💬 Integration Tip
Start with the default RapidOCR engine for quick setup without API keys, and configure environment variables for seamless switching to the cloud engine when higher accuracy is needed.
Scored Jun 19, 2026
Data analysis tool for Excel, CSV, Word, PDF, TXT, Markdown files. Use when user needs to analyze, summarize, or compare data from multiple files. Supports f...
PDF智能处理工具 v2.1 | PDF Smart Tool. 支持PDF转换、OCR识别、合并拆分、数字签名、批量处理、水印添加、加密解密。触发词:PDF、转换、识别。
Generate hand-drawn style diagrams, flowcharts, and architecture diagrams as PNG images from Excalidraw JSON
Convert public web pages into clean Markdown with markdown.new for AI workflows. Use when tasks require URL-to-Markdown conversion for summarization, RAG ing...
PDF扫描件转Word文档。支持中文OCR识别,自动裁掉页眉页脚,保留插图,彩色章节封面页保留为图片。使用百度OCR API(免费额度1000次/月)。当用户要求把扫描PDF转成文字/Word时触发。
【版权:青岛火一五信息科技有限公司 账号:huo15】Word/文档生成首选技能。触发词:写word、写文档、写个文档、重新写、重新生成、生成word、生成文档、创建word、创建文档、导出word、导出文档、下载word、下载文档、.docx、word文档、Word文档、写合同、写报价单、写说明书、写会议纪要、...