mineru-pdf-extractorExtract PDF content to Markdown using MinerU API. Supports formulas, tables, OCR. Provides both local file and online URL parsing methods.
Install via ClawdBot CLI:
clawdbot install A-I-R/mineru-pdf-extractorGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Accesses sensitive credential files or environment variables
/etc/passwdSends data to undocumented external endpoint (potential exfiltration)
UPLOAD → https://mineru.oss-cn-shanghai.aliyuncs.com/...Accesses system directories or attempts privilege escalation
/etc/cronCalls external URL not in known-safe list
https://mineru.net/Generated Mar 21, 2026
Researchers can extract content from PDF papers, including formulas and tables, to convert them into structured Markdown for literature reviews or data extraction. This aids in summarizing findings and integrating references into research databases efficiently.
Law firms and legal departments use this skill to parse contracts, case files, and regulations from PDFs into Markdown, enabling easier search, annotation, and compliance tracking. OCR support ensures accurate text extraction from scanned documents.
Financial analysts extract tables and textual data from PDF reports, such as earnings statements or market analyses, to automate data entry into spreadsheets or analytical tools. This streamlines financial modeling and reporting workflows.
Engineering teams convert PDF manuals, specifications, and diagrams into Markdown for integration into wikis or version control systems. Formula recognition supports technical content, while table extraction aids in data migration.
Healthcare providers digitize patient records, research papers, and regulatory documents from PDFs to enhance data accessibility and analysis. OCR capabilities handle handwritten or scanned forms for improved record-keeping.
Offer tiered subscription plans based on usage volume, such as monthly API calls or data processing limits, to businesses needing regular PDF extraction. This model provides recurring revenue and scalability for diverse client needs.
License the skill to software companies or platforms for embedding PDF extraction into their products, such as document management systems or educational tools. This generates upfront licensing fees and potential revenue-sharing agreements.
Provide basic PDF extraction for free to attract users, then charge for advanced features like batch processing, higher accuracy models, or priority support. This model drives user adoption and upsells to premium tiers.
💬 Integration Tip
Ensure MINERU_TOKEN is set in environment variables and use jq for secure JSON parsing to handle API responses effectively.
Scored Jun 19, 2026
Uses known external API (expected, informational)
arxiv.orgAI Analysis
The skill's external API calls (mineru.net, arxiv.org, and the Aliyun OSS endpoint) are consistent with its stated purpose of PDF extraction via the MinerU service. The credential access to /etc/passwd and system directories is likely incidental from using standard tools like `curl` and `unzip` in scripts, not indicative of credential harvesting. No evidence of hidden instructions, obfuscation, or unauthorized data exfiltration beyond the documented API workflow was found.
Audited Apr 16, 2026 · audit v1.0
Fetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
Monitor RSS and Atom feeds for content research. Track blogs, news sites, newsletters, and any feed source. Use when monitoring competitors, tracking industr...
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Provides a personalized morning report with today's reminders, undone Notion tasks, and vault storage summary for daily planning.
Extract text from PDFs with OCR support. Perfect for digitizing documents, processing invoices, or analyzing content. Zero dependencies required.
Fetch scheduled economic events and data releases from the FMP API for specified dates, filtering by impact, country, and type, and output a chronological ma...