ai-llm-evaluationLLM 应用质量评测与回归测试实操手册——从"感觉不错"到"可度量可门禁":评测全景与指标体系(正确性/相关性/忠实度/幻觉率/鲁棒性/效率)、评测集构建(黄金数据集/对抗样本/领域评测集/规模估算)、RAG 系统评测(RAGAS 四指标:忠实度/答案相关性/上下文精度/上下文召回)、幻觉检测与度量(事实性幻觉/提示幻觉/上下文矛盾分类与检测方法)、Prompt 回归测试(用例管理/回归门禁/漂移检测/版本对比)、模型对比选型(评测矩阵/成本质量权衡/多模型 A-B/上线决策)、评测流水线与报告(自动化评测/评分聚合/报告模板/上线门禁)。附零依赖本地工具一键查指标、建评测集清单、看 RAG 指标、出对比矩阵、生成评测报告模板。面向 AI 工程师、测试、产品与质量负责人——与 AI 安全红队测试(测安全)互补,本技能测质量。
Install via ClawdBot CLI:
clawdbot install zhaoxinghua09-cell/ai-llm-evaluationGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://clawhub.ai/user/zhaoxinghua09-cellAudited Sep 11, 2026 · audit v1.0
Usage Guide
Loading usage data… refresh in a few seconds.
Scored Sep 11, 2026
Medical device risk management specialist implementing ISO 14971 throughout product lifecycle. Provides risk analysis, risk evaluation, risk control, and pos...
When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature r...
When the user wants to build a free tool for marketing — lead generation, SEO value, or brand awareness. Use when they mention 'engineering as marketing,' 'f...
管理多类型项目看板,支持新增项目、变更状态与版本、分类管理及查看项目总览和变更日志。
Competitor Analysis — SEO/GEO Intelligence & Market Positioning. Analyze competitor SEO rankings, AI search citations, content strategy, and market posi...
Conduct structured PHQ-9 depression symptom screening and submit the completed assessment for evaluation.