mjk39966-document-parserParse and extract content from .docx, .pdf, and .txt documents. Extracts plain text and tables for analysis. Use when the user uploads a document file or ask...
Install via ClawdBot CLI:
clawdbot install mjk39966-glitch/mjk39966-document-parserGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://www.python.org/downloads/Audited Apr 18, 2026 · audit v1.0
Generated Oct 6, 2026
A financial analyst receives quarterly earnings reports and annual 10-K filings in PDF and DOCX formats. Using the document-parser skill, they extract revenue tables and text sections to quickly calculate key metrics like YoY growth and profit margins. This accelerates due diligence and enables faster investment recommendations.
A paralegal uploads multiple contract documents in DOCX format and needs to locate specific clauses such as indemnification or termination conditions. The skill extracts all text and tables, allowing the legal team to search and analyze terms efficiently. This reduces manual review time and minimizes oversight risks.
A graduate student collects research papers in PDF format and wants to extract methodologies, results tables, and key findings. The skill parses each document and provides structured JSON output for further analysis or summarization. This streamlines literature reviews and data synthesis.
An insurance adjuster receives claim forms, policy documents, and receipts in various formats (PDF, DOCX, TXT). The skill extracts relevant text and tabular data to validate coverage and calculate payouts. This automates data entry and speeds up claim resolution.
A product manager gathers competitor brochures, whitepapers, and pricing sheets in PDF and DOCX. The skill extracts text and tables to compare features and pricing across competitors. This informs product strategy and positioning.
Offer a REST API that accepts document uploads and returns parsed text and tables in JSON. Pricing is based on the number of pages processed or API calls. This model targets developers and enterprises needing programmatic document extraction.
A web application where users can upload documents, view extracted content, and perform analyses like table calculations or keyword search. Subscription tiers are based on storage, processing volume, and advanced features. This serves business users who need an intuitive interface.
License the document parsing technology as a white-label component to be embedded into existing enterprise software (e.g., CRM, ERP, legal practice management). The model includes custom integration and support contracts. Revenue comes from annual licensing fees.
💬 Integration Tip
Ensure all dependencies are installed by running the provided installation scripts before first use. For large PDFs, consider processing asynchronously and monitor resource usage.
Scored Oct 6, 2026
PDF智能处理工具 v2.1 | PDF Smart Tool. 支持PDF转换、OCR识别、合并拆分、数字签名、批量处理、水印添加、加密解密。触发词:PDF、转换、识别。
Generate hand-drawn style diagrams, flowcharts, and architecture diagrams as PNG images from Excalidraw JSON
Convert public web pages into clean Markdown with markdown.new for AI workflows. Use when tasks require URL-to-Markdown conversion for summarization, RAG ing...
PDF扫描件转Word文档。支持中文OCR识别,自动裁掉页眉页脚,保留插图,彩色章节封面页保留为图片。使用百度OCR API(免费额度1000次/月)。当用户要求把扫描PDF转成文字/Word时触发。
Word文档处理工具套件,提供Word文档的创建、读取、内容提取和基本处理功能。
Microsoft SharePoint and OneDrive integration with managed OAuth. Manage sites, lists, libraries, files, folders, permissions, content types, and SharePoint...