skill-evaluationEvaluate any AI skill's quality through step-by-step diagnosis — measuring trigger accuracy, per-step execution (completion/correctness/quality), efficiency,...
Install via ClawdBot CLI:
clawdbot install rivin-dong/skill-evaluationGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
eval(Calls external URL not in known-safe list
https://clawhub.ai/user/rivin-dongAudited Jun 4, 2026 · audit v1.0
Generated Jul 28, 2026
A company develops a skill to handle customer refund requests. This evaluation diagnoses whether the skill triggers correctly on refund intent, executes each step with high correctness and quality, and flags any bad cases like incorrect refund amounts or missing escalation steps, enabling iterative optimization before production deployment.
A marketing agency wants to improve an AI skill that generates ad copy. The evaluation measures trigger accuracy for copywriting requests, checks per-step execution for tone consistency and brand compliance, and identifies bad cases such as off-brand messaging, providing actionable fixes for the prompt.
A health tech startup needs to test a skill that suggests possible diagnoses based on symptoms. The evaluation ensures the skill triggers only for medical queries, executes with high correctness (avoiding false positives), and generates a report highlighting safety concerns or missing disclaimers.
An e-commerce platform is iterating on a product recommendation skill. The evaluation runs on two versions (v1 and v2), comparing trigger precision, step-by-step execution quality, and efficiency, pinpointing regressions and guiding optimization.
A software company wants to deploy a skill that generates code snippets. The evaluation tests trigger accuracy for coding queries, checks correctness and quality of generated code, and identifies bad cases like syntax errors or insecure patterns, ensuring the skill meets production standards.
Offer tiered monthly subscriptions for skill evaluation. Basic tier includes quick evaluation with 4 test cases, premium tier includes deep evaluation with up to 12 test cases and iterative optimization support, generating recurring revenue.
Charge per evaluation run, with optional credits for optimization reports. Each run costs a fixed fee, and users purchase optimization credits for detailed root cause analysis and skill fixes, enabling high-volume usage without long-term commitment.
License the evaluation pipeline to enterprises for integration into their CI/CD workflow. Provide API access, multi-skill evaluation support, and priority support for a flat annual fee, appealing to companies with frequent skill updates.
💬 Integration Tip
Integrate the evaluation pipeline into your CI/CD loop by triggering automated evaluations on each skill version commit, using the generated summary.md to gate deployments until all bad cases are resolved.
Scored Jun 4, 2026
Humanize AI-generated text to bypass detection. This humanizer rewrites ChatGPT, Claude, and GPT content to sound natural and pass AI detectors like GPTZero,...
Generate ideas fast. Adapt depth and structure to what the user actually needs.
通过调用 Prana 平台上的远程 agent 完成以下处理:基于100个热门TradingView Pine Script指标转换的Python技术分析工具集,提供专业的技术指标计算、分析和可视化功能 IMPORTANT: This skill has a mandatory step-by-step proc...
Provides a structured screening for stress perception using the PSS-10 scale as an independent skill in ClawHub.
Structured self-improvement system for AI agents with confidence decay, cross-agent sharing, and anomaly detection. Use when: (1) After debugging sessions to...
Spawns real AI-powered OpenClaw sub-sessions to run multiple specialized agents concurrently for content, dev, QA, docs, and autonomous workflows.