yuyonghao-evaluation-suiteProvides API for evaluating RAG quality, logical reasoning, and detecting hallucinations in AI-generated content with batch support.
Install via ClawdBot CLI:
clawdbot install yuyonghao-123/yuyonghao-evaluation-suiteGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Generated May 9, 2026
Use RAG evaluation to continuously assess the quality of answers generated by a customer support chatbot. Ensure retrieved documents are relevant and answers are accurate, reducing response errors.
Apply hallucination detection to verify AI-generated legal documents against provided context. Prevent fabricated clauses or citations that could lead to legal risks.
Evaluate the logical reasoning steps of an AI tutor when solving math or science problems. Ensure the tutor's explanations are coherent and lead to correct answers.
Use the evaluation suite to test AI systems that provide diagnostic suggestions based on medical literature. Check for hallucinated symptoms or incorrect reasoning that could compromise patient safety.
Assess AI-generated product descriptions for factual accuracy against retrieved product specs. Detect hallucinations like false features or pricing to maintain trust.
Offer the evaluation suite as a cloud-based service where clients pay per evaluation or via monthly subscription. Integrate with CI/CD pipelines for automated testing.
Provide tailored evaluation frameworks for enterprise AI deployments. Includes threshold tuning, custom metrics, and integration support.
Release core evaluation capabilities as open source to drive adoption, while charging for advanced features like batch processing, reporting dashboards, and priority support.
💬 Integration Tip
Start by integrating the evaluator into your CI/CD pipeline to run evaluations on every model update or before deployment. Use the threshold parameters to align with your quality standards.
Scored May 9, 2026
Humanize AI-generated text to bypass detection. This humanizer rewrites ChatGPT, Claude, and GPT content to sound natural and pass AI detectors like GPTZero,...
AI brainstorming and strategy thinking partner powered by CellCog. Reasoning, problem-solving, ideation, strategic planning — then execution across every modality: research, documents, visuals, data, prototypes. Think, build, review, repeat.
Generate ideas fast. Adapt depth and structure to what the user actually needs.
Evaluate any AI skill's quality through step-by-step diagnosis — measuring trigger accuracy, per-step execution (completion/correctness/quality), efficiency,...
通过调用 Prana 平台上的远程 agent 完成以下处理:基于100个热门TradingView Pine Script指标转换的Python技术分析工具集,提供专业的技术指标计算、分析和可视化功能 IMPORTANT: This skill has a mandatory step-by-step proc...
Provides a structured screening for stress perception using the PSS-10 scale as an independent skill in ClawHub.