novel-scraper-pro智能小说抓取工具 V6,支持自动翻页、分页补全、章节号自动解析、内存监控、中断续抓。 使用 curl+BeautifulSoup 抓取笔趣阁等小说网站,输出格式化 TXT 文件。 默认每 10 章合并为一个文档,避免文件零散分布。 自动检测分页并补全,智能跳过非小说内容(作者感言、抽奖预告等)。 内存监控和中断续...
Install via ClawdBot CLI:
clawdbot install yuzhihui886/novel-scraper-proGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
rm -rf /Accesses system directories or attempts privilege escalation
/proc/Calls external URL not in known-safe list
https://www.bqquge.com/4/1962AI Analysis
The skill's primary function is web scraping of public novel websites, which aligns with its stated purpose. While it accesses external URLs not on a pre-approved list, this is inherent to its scraping functionality and doesn't indicate credential harvesting or data exfiltration. The 'UNSAFE_SHELL' and 'SUSPICIOUS_PERMISSIONS' signals appear to be from hypothetical examples in documentation rather than actual malicious code in the provided skill definition.
Generated Sep 26, 2026
Individual readers who want to build offline TXT libraries of web novels from sites like Biquge can batch-scrape hundreds of chapters with automatic pagination completion and resume support. The merge-interval feature groups content into digestible 10-chapter files, making it easy to read on e-readers or mobile devices without internet access.
Translation studios or freelancers scraping source material for localization can use the scraper to reliably extract clean chapter text, skipping author notes and giveaway announcements that pollute translation output. The chapter-number auto-parsing and continuation checking ensure no chapters are missed during large-scale extraction projects.
Publishing houses and literary agencies monitoring popular web novels can archive competitor content for market analysis, trend research, and IP acquisition evaluation. The cache system and interruption resume make it practical to maintain large, continuously updated archives of serialized fiction across multiple titles.
Researchers building Chinese-language NLP models or recommendation systems can use the tool to curate structured novel datasets, leveraging strict quality validation to filter low-quality chapters. Output as formatted TXT files with consistent naming conventions simplifies downstream tokenization, indexing, and annotation workflows.
Community operators running fan forums or reading groups for specific novel genres can scrape and redistribute chapter bundles to members, especially in regions with unreliable internet access. The SPA-force flag extends coverage to modern JavaScript-rendered novel sites that traditional scrapers cannot handle.
Offer a hosted version of Novel Scraper Pro with a free tier limited to single-chapter or small-batch scraping, and paid tiers unlocking bulk chapter ranges, priority queues, and cloud storage of scraped TXT files. This removes the Python setup barrier for non-technical readers who just want their novels downloaded.
Aggregate and clean large corpora of Chinese web novels using the scraper, then license these structured datasets to LLM training companies, translation AI firms, and recommendation engine developers. The chapter-level caching, quality validation, and consistent formatting create higher-value data than raw scraped HTML.
Provide a white-label service where publishing clients submit novel URLs and receive regularly updated, complete TXT archives with quality reports and missing-chapter alerts. The resume-after-interruption and memory monitoring features make long-running, unattended scraping reliable enough for enterprise SLAs.
💬 Integration Tip
Install via clawhub and run with Python 3.8+ and beautifulsoup4; test with a single --url first, then scale to --chapters ranges while relying on default resume and memory monitoring for long jobs. Clear /tmp/novel_scraper_cache/ and progress.json when you need a fresh re-scrape of previously fetched chapters.
Scored Jun 3, 2026
Audited Apr 17, 2026 · audit v1.0
A fast Rust-based headless browser automation CLI with Node.js fallback that enables AI agents to navigate, click, type, and snapshot pages via structured commands.
Uses a headless browser to navigate web pages, interact with elements, and extract clean, readable text content from URLs.
Headless browser automation CLI optimized for AI agents with accessibility tree snapshots and ref-based element selection
High-performance browser automation for heavy scraping, multi-tab management, and precise DOM extraction. Use this when you need speed, reliability, or advanced state management (cookies/local storage) beyond standard web fetching.
Ultimate stealth browser automation with anti-detection, Cloudflare bypass, CAPTCHA solving, persistent sessions, and silent operation. Use for any web automation requiring bot detection evasion, login persistence, headless browsing, or bypassing security measures. Triggers on "bypass cloudflare", "solve captcha", "stealth browse", "silent automation", "persistent login", "anti-detection", or any task needing undetectable browser automation. When user asks to "login to X website", automatically use headed mode for login, then save session for future headless reuse.
Browser automation CLI (browser-act) for AI agents. MUST trigger when: (1) user mentions 'browser-act' in any form, or user needs to: (2) open/visit/browse/c...