openclaw-skill-evalSkill evaluation framework. Use when: testing trigger rate, quality compare (with/without skill), or model comparison. Runs via sessions_spawn + sessions_his...
Install via ClawdBot CLI:
clawdbot install xiaoxing9/openclaw-skill-evalGrade Good — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Potentially destructive shell commands in tool definitions
eval(Calls external URL not in known-safe list
https://api.zephyr.internalAI Analysis
The skill is an evaluation framework for testing other skills and operates primarily on local files and OpenClaw's session API. The external URL 'https://api.zephyr.internal' is part of a bundled, fictional test skill for internal validation, not a real external call. No evidence of credential harvesting, data exfiltration, or hidden malicious instructions exists.
Audited Apr 17, 2026 · audit v1.0
Generated May 21, 2026
Run trigger rate evaluation to measure how often AI agents correctly invoke a skill's SKILL.md description in relevant contexts. Includes positive and negative test cases to compute recall, specificity, precision, and F1 score, helping improve skill description accuracy.
Compare AI agent output quality with and without a skill enabled using assertion-based tests. Determine whether the skill improves relevance, accuracy, or completeness for benchmark queries.
Evaluate quality and speed of different AI models (e.g., Haiku, Sonnet, Opus) on the same skill to determine the cost-performance sweet spot before deployment. Run automated tests across models and compare results.
Analyze false negatives and false positives from trigger rate tests to identify gaps in skill descriptions. Get actionable recommendations for rewording or restructuring SKILL.md to improve invocation accuracy.
Test trigger rate detection using a bundled fictional 'Zephyr API' skill to validate that the evaluation framework works correctly. Requires manual installation of the test skill to simulate a skill unknown to the AI model.
Offer a cloud-based service where teams can upload skill definitions and run automated evaluations on demand. Revenue from subscription tiers based on evaluation volume and storage.
Provide expert analysis of evaluation results, recommending improvements to skill descriptions and triggers for clients. Revenue from hourly consulting or project-based engagements.
License the evaluation framework to enterprises for internal use, including customization and integration support. Revenue from annual licensing fees with premium support tiers.
💬 Integration Tip
Ensure your environment has sessions_spawn and sessions_history enabled, and that the skill you want to evaluate is accessible via OpenClaw's extraDirs. Start with the quick eval command to test the workflow.
Scored Jun 29, 2026
API reference for CoinMarketCap DEX endpoints including token lookup, pools, transactions, trending, and security analysis. Use this skill whenever the user...
Use when designing REST or GraphQL APIs, creating OpenAPI specifications, or planning API architecture. Invoke for resource modeling, versioning strategies, pagination patterns, error handling standards.
ServiceNow IT Service Management integration with API key authentication. Manage incidents, problems, change requests, tasks, CMDB records, service catalog i...
This skill should be used when writing documentation for codebases, including README files, architecture documentation, code comments, and API documentation. Use this skill when users request help documenting their code, creating getting-started guides, explaining project structure, or making codebases more accessible to new developers. The skill provides templates, best practices, and structured approaches for creating clear, beginner-friendly documentation.
API Design Reviewer
Cut AI token costs by 60-85% with deterministic query validation and intelligent caching. Standalone or integrates with Company Brain Core OS. Free, open-sou...