skill-evalscopeTranslates natural language requests into evalscope CLI commands. Core capabilities: (1) Model accuracy evaluation (eval) — runs 156+ benchmarks (Math, Codin...
Install via ClawdBot CLI:
clawdbot install yunnglin/skill-evalscopeGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
http://localhost:8000/v1/chat/completionsUses known external API (expected, informational)
api.openai.comAudited Apr 17, 2026 · audit v1.0
Generated May 7, 2026
A data scientist has trained a new model checkpoint and wants to evaluate its accuracy on math and coding benchmarks. Using EvalScope, they can run 'evalscope eval --model <path> --datasets gsm8k humaneval' to quickly get scores and compare with baseline models.
An MLOps engineer needs to measure throughput and latency of a deployed model API under varying loads. They use 'evalscope perf --model my-model --api-url <url> --concurrency 1,2,4,8' to generate a throughput-latency report and identify bottlenecks.
A researcher wants to find all benchmarks related to 'reasoning' for a survey. They run 'evalscope benchmark-info --list --tag reasoning' to retrieve a curated list with metadata and sample examples, then export the results.
A team lead wants to visualize and compare evaluation results from multiple model runs. They launch the Gradio UI with 'evalscope app' in the output directory, then interactively filter and compare metrics across experiments.
A QA engineer wants to test Claude's performance on MMLU before deployment. They configure 'evalscope eval --model claude-3-5-sonnet --eval-type anthropic_api --datasets mmlu --api-key sk-ant-xxx' and run a limited sample to validate integration.
Offer basic evaluation (e.g., free tier with limited benchmarks and concurrency) and premium tiers for high-throughput performance testing and priority support. Monetize via monthly subscriptions.
Provide EvalScope as a managed API where customers pay per evaluation run or per sample. Charge per dataset-model pair with volume discounts, targeting startups that don't want to manage infrastructure.
Deploy EvalScope as an on-premise or private cloud solution for enterprises with stringent data security. Include custom benchmark integration, SLAs, and dedicated support. Revenue from annual contracts.
💬 Integration Tip
Start with mock_llm mode (--eval-type mock_llm) to validate your workflow without a real model, then swap to your actual model source.
Scored Jun 29, 2026
Use this skill when users need to search academic papers, download research documents, extract citations, or gather scholarly information. Triggers include: requests to "find papers on", "search research about", "download academic articles", "get citations for", or any request involving academic databases like arXiv, PubMed, Semantic Scholar, or Google Scholar. Also use for literature reviews, bibliography generation, and research discovery. Requires OpenClawCLI installation from clawhub.ai.
Get institutional-grade CEO performance analytics for S&P 500 companies. Proprietary scores: CEORaterScore (composite), AlphaScore (market outperformance), R...
Creates formal academic research papers following IEEE/ACM formatting standards with proper structure, citations, and scholarly writing style. Use when the user asks to write a research paper, academic paper, or conference paper on any topic.
Search, download, and summarize academic papers from arXiv. Built for AI/ML researchers.
Suggest optimal academic venues for paper submission based on paper content, novelty level, and author's goals. Activated when Chopin asks 'which journal sho...
Baidu Scholar Search - Search Chinese and English academic literature (journals, conferences, papers, etc.)