advanced-evaluationThis skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", o...
Install via ClawdBot CLI:
clawdbot install karmaent/advanced-evaluationGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://eugeneyan.com/writing/llm-evaluators/Uses known external API (expected, informational)
arxiv.orgAudited Apr 17, 2026 · audit v1.0
Generated May 23, 2026
Automate grading of AI assistant responses for factual accuracy, instruction following, and safety. Use direct scoring with calibrated rubrics to flag poor-quality outputs.
Compare two versions of a conversational AI on subjective criteria like tone and persuasiveness using pairwise comparison with position bias mitigation.
Implement an LLM-as-judge system to score student essays on clarity, argument strength, and evidence use, reducing human grading workload.
Detect bias in automated content moderation by analyzing disagreement patterns between LLM judges and human reviewers, focusing on systematic errors.
Build a production-grade pipeline that continuously evaluates retrieval-augmented generation outputs using direct scoring for answer relevance and factual accuracy.
Offer API-based LLM evaluation services for companies building AI applications. Charge per evaluation call with tiered pricing based on volume and rubric complexity.
Provide a dashboard for teams to design rubrics, run evaluations, and track metric trends. Monthly or annual subscription with seats and evaluation credit tiers.
Perform in-depth bias analysis on client evaluation systems using the bias mitigation protocols. Deliver reports with remediation recommendations.
💬 Integration Tip
Start by implementing direct scoring for one objective criterion (e.g., factual accuracy) with a 1-3 scale, then layer in pairwise comparison once you have reliable scoring.
Scored May 23, 2026
Medical device risk management specialist implementing ISO 14971 throughout product lifecycle. Provides risk analysis, risk evaluation, risk control, and pos...
When the user wants to plan a product launch, feature announcement, or release strategy. Also use when the user mentions 'launch,' 'Product Hunt,' 'feature r...
AI project management powered by CellCog. Knowledge workspaces, document upload, AI-processed context trees, signed URL retrieval. Works standalone or as CellCog chat context.
When the user wants to build a free tool for marketing — lead generation, SEO value, or brand awareness. Use when they mention 'engineering as marketing,' 'f...
管理多类型项目看板,支持新增项目、变更状态与版本、分类管理及查看项目总览和变更日志。
Competitor Analysis — SEO/GEO Intelligence & Market Positioning. Analyze competitor SEO rankings, AI search citations, content strategy, and market posi...