xh-smart-scraper智能网页数据采集器。自动识别网页结构,批量抓取列表/表格/详情页数据,支持导出JSON/CSV/Excel。内置反爬策略适配。
Install via ClawdBot CLI:
clawdbot install cjstate/xh-smart-scraperGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://example.com/productsAudited Apr 18, 2026 · audit v1.0
Generated May 9, 2026
自动抓取电商平台商品列表和详情页,提取标题、价格、图片等字段,支持分页采集。用于实时监控竞品价格变动,辅助定价决策。
采集房产网站上的房源信息,如房价、面积、位置等,支持配置字段选择。可用于生成市场报告或建立房源数据库。
从招聘网站抓取职位列表,提取职位名称、公司、薪资等字段,支持设置请求延迟以避免封禁。用于人才市场分析或内部职位库建设。
自动识别学术网站上的论文列表和详情页,提取标题、作者、摘要等,导出为JSON或CSV。用于文献综述或研究分析。
从社交媒体或论坛批量抓取用户评论和帖子内容,支持反爬机制。用于舆情监控或情感分析。
将智能网页爬虫包装为云端服务,按数据量或API调用次数收费。用户无需安装,直接通过Web界面配置采集任务。
为B端客户提供定制化的网页数据采集方案,包括复杂网站的数据抓取、清洗和定期更新。
开源基础版吸引用户,企业版提供高级功能如代理池、数据库直存、技术支持等。
💬 Integration Tip
只需安装Node.js和npm,通过命令行或配置文件即可快速启动采集,输出结果可直接导入数据分析工具或数据库。
Scored May 9, 2026
A fast Rust-based headless browser automation CLI with Node.js fallback that enables AI agents to navigate, click, type, and snapshot pages via structured commands.
Uses a headless browser to navigate web pages, interact with elements, and extract clean, readable text content from URLs.
Headless browser automation CLI optimized for AI agents with accessibility tree snapshots and ref-based element selection
通过已登录的Edge或Chrome浏览器,利用Chrome DevTools Protocol执行JS渲染页面的导航、点击、截图及数据提取等自动化操作。
Send and receive SMS/RCS via Google Messages web interface (messages.google.com). Use when asked to "send a text", "check texts", "SMS", "text message", "Google Messages", or forward incoming texts to other channels.
High-performance browser automation for heavy scraping, multi-tab management, and precise DOM extraction. Use this when you need speed, reliability, or advanced state management (cookies/local storage) beyond standard web fetching.