pdf-utilsPDF Utils enables OCR of image-based PDFs, extraction of arXiv IDs from text or OCR output, and scriptable PDF tasks like merging, splitting, and rendering.
Install via ClawdBot CLI:
clawdbot install wangwllu/pdf-utilsGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://img.shields.io/badge/license-MIT-blue.svgUses known external API (expected, informational)
arxiv.orgAudited Apr 18, 2026 · audit v1.0
Generated Oct 6, 2026
Researchers processing downloaded arXiv papers need to extract all cited arXiv IDs from text-based PDFs and optionally batch-download referenced papers into a local folder. The skill's extract_refs.py automates what would otherwise be hours of manual reference hunting.
Libraries and archives converting image-based scanned documents into searchable text using Tesseract OCR via ocr_pdf.py. Combined with --extract-refs, staff can mine citations from century-old scanned academic journals.
Law firms processing scanned discovery PDFs need OCR plus repeatable merge/split operations across thousands of filings. pdf_ops.py handles batching while ocr_pdf.py makes image-only exhibits text-searchable.
A newsroom or newsletter operator ingests arXiv PDFs, extracts referenced paper IDs, downloads them, and assembles curated weekly digests for subscribers. The pipeline runs entirely scriptable via extract_refs.py with --download.
R&D teams build internal knowledge bases from scanned patent filings and technical papers, OCR'ing them and mining cross-references. Page-range OCR keeps memory usage sane on multi-hundred-page filings.
Offer a subscription platform that ingests user-uploaded PDFs and returns OCR text, extracted arXiv IDs, and auto-downloaded references. Wraps the skill's scripts behind a web API with queueing for large batch jobs.
Agency service that OCRs and structures scanned document collections for law firms, libraries, and archives. Uses pdf_ops.py and ocr_pdf.py as the production pipeline, charging per page or per project.
Sell a desktop or VSCode/Jupyter plugin to researchers that integrates arXiv mining, reference downloading, and PDF merging directly into their workflow. Monetized through one-time licenses or institutional site licenses.
💬 Integration Tip
Install the Tesseract/PyMuPDF dependencies first and route summarization or Q&A tasks to the built-in pdf tool instead of this skill; reserve these scripts for OCR, arXiv extraction, and repeatable local PDF ops.
Scored Oct 6, 2026
Fetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
Monitor RSS and Atom feeds for content research. Track blogs, news sites, newsletters, and any feed source. Use when monitoring competitors, tracking industr...
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Provides a personalized morning report with today's reminders, undone Notion tasks, and vault storage summary for daily planning.
Extract text from PDFs with OCR support. Perfect for digitizing documents, processing invoices, or analyzing content. Zero dependencies required.
Fetch scheduled economic events and data releases from the FMP API for specified dates, filtering by impact, country, and type, and output a chronological ma...