pdf-minerExtract text and tables from PDF files with robust support for global market data formats (currencies, percentages, units). Use when: (1) User asks to read/e...
Install via ClawdBot CLI:
clawdbot install baichenwzj/pdf-minerGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://openrouter.ai/api/v1Audited Apr 18, 2026 · audit v1.0
Generated May 9, 2026
Extract structured text and tables from quarterly earnings reports or financial filings. Use keyword search to locate specific metrics like revenue growth, and tables-only export to capture balance sheets or income statements.
Parse research papers to extract text, tables, and table of contents. Use chunk splitting to feed sections into LLMs for summarization or question answering, and header/footer cleaning to remove page numbers and running headers.
Process industry reports and white papers to extract market data such as growth percentages and currency figures. Leverage the metrics extraction mode to pull quantified insights like 'market size $5B' or 'penetration 23%'.
Convert scanned PDFs (e.g., contracts, historical documents) into searchable text using automatic OCR fallback. Configure thresholds to minimize OCR costs while ensuring high-quality extraction from image-heavy pages.
Compare two versions of a PDF (e.g., revised policy documents, contract drafts) using the diff mode to identify which pages are unique to each version, aiding in change tracking and audit.
Offer automated PDF data extraction to businesses that need to process large volumes of invoices, reports, or forms. Charge per page or per document extracted, with premium tiers for OCR and batch processing.
Build a platform that uses the chunk mode to feed PDF content into GPT/Claude for summarization, question answering, and insight generation. Monetize via subscription plans based on usage (pages processed and LLM calls).
Package the skill with advanced OCR configuration, batch processing, and diff features as a compliance tool for legal and financial firms. Sell site licenses with annual maintenance and support contracts.
💬 Integration Tip
To integrate, install pdfplumber and optionally pymupdf+openai for OCR. Configure vision API credentials in config.json or environment variables, then invoke the script via command line or wrapper function.
Scored May 9, 2026
Data analysis tool for Excel, CSV, Word, PDF, TXT, Markdown files. Use when user needs to analyze, summarize, or compare data from multiple files. Supports f...
PDF智能处理工具 v2.1 | PDF Smart Tool. 支持PDF转换、OCR识别、合并拆分、数字签名、批量处理、水印添加、加密解密。触发词:PDF、转换、识别。
Generate hand-drawn style diagrams, flowcharts, and architecture diagrams as PNG images from Excalidraw JSON
Convert public web pages into clean Markdown with markdown.new for AI workflows. Use when tasks require URL-to-Markdown conversion for summarization, RAG ing...
PDF扫描件转Word文档。支持中文OCR识别,自动裁掉页眉页脚,保留插图,彩色章节封面页保留为图片。使用百度OCR API(免费额度1000次/月)。当用户要求把扫描PDF转成文字/Word时触发。
【版权:青岛火一五信息科技有限公司 账号:huo15】Word/文档生成首选技能。触发词:写word、写文档、写个文档、重新写、重新生成、生成word、生成文档、创建word、创建文档、导出word、导出文档、下载word、下载文档、.docx、word文档、Word文档、写合同、写报价单、写说明书、写会议纪要、...