baidu-doc-pipeline-parser调用百度文档解析API解析文档。支持PDF、Word、Excel、PPT、图片等18+格式。提取文本、表格、版面分析、OCR识别及RAG文档分块。当用户需要解析文档、提取文本/表格、分析文档结构、处理扫描件时使用。触发词:文档解析、PDF解析、Word解析、表格提取、OCR、文档分析、提取文本、文档结构、扫描识别。
Install via ClawdBot CLI:
clawdbot install maglanyulan/baidu-doc-pipeline-parserGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Hardcoded API key or token pattern found in skill definition
sk-3zy9Bg8CH...Calls external URL not in known-safe list
https://ai.baidu.com/ai-doc/OCR/dk3iqnq51Uses known external API (expected, informational)
aip.baidubce.comAI Analysis
The skill uses a legitimate external API (Baidu Document AI) for document parsing, which is consistent with its stated purpose. The credential pattern detected is likely a false positive from example code or documentation, not an actual hardcoded key. No hidden instructions, obfuscation, or data exfiltration mechanisms were found.
Generated Oct 6, 2026
法务和财务团队批量上传PDF、Word、扫描件合同与发票,自动提取正文文本、表格金额和版面结构,将非结构化文档转换为可检索的结构化数据,用于归档与比对。
开发者将产品手册、政策文件、研究报告解析为带语义切分的chunks,并获取Markdown格式内容,直接灌入向量数据库构建企业问答知识库RAG应用。
研究机构批量解析中英日韩多语言PDF论文,抽取公式、图表位置、标题层级和参考文献结构,助力文献综述、专利检索和知识图谱构建。
银行、券商、审计机构解析Excel财报、PPT路演材料和扫描版审计底稿,识别表格数据与页面类型,实现数据自动抽取和风险合规审查。
政府、医院、档案馆将历史纸质档案扫描后调用OCR识别,配合角度矫正和版面分析,将图片型文档转化为可全文检索的电子档案。
面向中小企业提供网页端文档解析工作台,按月/年订阅,支持批量上传、解析历史查询和Markdown导出,无需自建API对接。
将百度文档解析API封装为统一接口,附加文档分块、格式转换、任务队列管理等增值能力,按调用量向开发者或集成商计费。
为垂直行业客户交付端到端文档到知识库的方案,包含解析流水线、向量化、检索问答系统定制开发与运维。
💬 Integration Tip
该接口为异步任务模式,需先提交请求获取task_id再轮询查询结果,并注意结果链接仅30天有效;生产环境务必处理轮询超时、额度不足报错和50MB以上文件改走file_url的逻辑。
Scored Oct 6, 2026
Audited May 10, 2026 · audit v1.0
Fetch and read transcripts from YouTube videos. Use when you need to summarize a video, answer questions about its content, or extract information from it.
Monitor RSS and Atom feeds for content research. Track blogs, news sites, newsletters, and any feed source. Use when monitoring competitors, tracking industr...
用 MinerU API 解析 PDF/Word/PPT/图片为 Markdown,支持公式、表格、OCR。适用于论文解析、文档提取。
Provides a personalized morning report with today's reminders, undone Notion tasks, and vault storage summary for daily planning.
Extract text from PDFs with OCR support. Perfect for digitizing documents, processing invoices, or analyzing content. Zero dependencies required.
Fetch scheduled economic events and data releases from the FMP API for specified dates, filtering by impact, country, and type, and output a chronological ma...