llm-eval-routerShadow-test local Ollama models against a cloud baseline with a multi-judge ensemble. Automatically promotes models when statistically proven equivalent — re...
Install via ClawdBot CLI:
clawdbot install nissan/llm-eval-routerGrade Fair — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://github.com/reddinft/skill-llm-eval-routerAudited Apr 16, 2026 · audit v1.0
Generated Mar 1, 2026
A fintech startup uses the skill to evaluate local models for summarizing quarterly financial reports, comparing them against Claude as a baseline. After collecting 200+ runs, they promote a local model to handle routine summaries, reducing API costs by 80% while maintaining quality through continuous monitoring.
An e-commerce company employs the skill to test local models for classifying customer support tickets into categories like refunds or technical issues. They use the multi-judge ensemble to ensure accuracy, promoting a model after it proves equivalent to cloud models, cutting down on expensive API calls for high-volume ticket processing.
A legal tech firm uses the skill to evaluate local models for analyzing contract clauses, with ground truth from Claude. They apply per-task weight overrides for analyze tasks to prioritize semantic similarity, ensuring reliable promotion of models that match cloud quality for non-critical legal reviews, saving on API expenses.
A social media platform integrates the skill to test local models for filtering inappropriate content, using cloud models as a baseline. After statistical validation, they promote a local model to handle initial filtering, reducing latency and costs while maintaining safety through demotion triggers if quality drops.
A healthcare analytics company uses the skill to evaluate local models for extracting structured data from patient notes, with ground truth from Anthropic. They leverage the deterministic validators for every run and promote models after meeting the 0.95 mean score threshold, enabling cost-effective data processing while ensuring compliance through local inference.
Offer the skill as a cloud-based service with tiered pricing based on usage volume, providing automated model evaluation and routing for enterprises. Revenue comes from monthly subscriptions, with premium tiers including advanced analytics and custom task type configurations.
Sell licenses for on-premise deployment to organizations with strict data privacy requirements, such as government or healthcare. Revenue is generated through one-time license fees and annual support contracts, with optional add-ons for integration with existing AI infrastructure.
Provide consulting services to help companies implement and customize the skill for specific use cases, such as optimizing task weights or integrating with local Ollama models. Revenue comes from project-based fees and ongoing maintenance agreements, targeting businesses new to AI cost optimization.
💬 Integration Tip
Ensure Ollama is running with capable models and set up per-task weight overrides in config files to align with specific evaluation needs, such as reducing structural weight for analyze tasks.
Scored May 30, 2026
Use CodexBar CLI local cost usage to summarize per-model usage for Codex or Claude, including the current (most recent) model or a full model breakdown. Trigger when asked for model-level usage/cost data from codexbar, or when you need a scriptable per-model summary from codexbar cost JSON.
Check Antigravity account quotas for Claude and Gemini models. Shows remaining quota and reset times with ban detection.
使用豆包(火山引擎)语音合成大模型 API 将文本转换为语音音频文件。支持声音复刻音色(S_ 开头的音色ID)和官方预置音色。当用户要求"语音合成"、"文字转语音"、"TTS"、"朗读文本"、"生成语音"、"用我的声音读"、"豆包语音"、"声音复刻合成"等相关请求时,务必使用此 skill。即使用户只是说"帮我把...
Intelligent model routing for sub-agent task delegation. Choose the optimal model based on task complexity, cost, and capability requirements. Reduces costs...
自动生成科技新闻摘要。从多个来源(RSS、Twitter、GitHub、Web Search)抓取科技新闻,整合后生成摘要。
让 AI 代理根据对话内容自动选择最合适的模型。四层识别(系统过滤→关键词→指示词→语义相似度),四池架构(高速/智能/人文/代理),五分支路由,全自动 Fallback 回路。支持 trigger_groups_all 非连续词组命中。