llm-testerLLM 模型对比测试工具。支持多模型批量对比测试,自动记录耗时、Token 消耗、成功率,生成 JSON 格式对比报告。当需要评估不同 LLM 模型在特定任务上的表现时使用。
Install via ClawdBot CLI:
clawdbot install yuzhihui886/llm-testerGrade Limited — based on market validation, documentation quality, package completeness, maintenance status, and authenticity signals.
Calls external URL not in known-safe list
https://coding.dashscope.aliyuncs.com/v1/chat/completions`Audited Apr 17, 2026 · audit v1.0
Generated May 23, 2026
A company wants to choose between several LLMs for their customer service chatbot. They can use LLM Tester to benchmark models on response quality, latency, and cost using sample customer queries and prompt templates.
A marketing team needs to evaluate different prompt variations for generating product descriptions. LLM Tester can compare multiple models and prompts to find the best combination for quality and speed.
An AI writing tool provider wants to ensure new model versions maintain quality. Using LLM Tester with standardized test samples and prompts, they can automatically generate comparison reports on accuracy and consistency.
A startup needs to balance cost and performance when deploying LLMs. LLM Tester records token consumption and response times, enabling data-driven decisions on which model to use for different tasks.
Researchers studying language models can use LLM Tester to run controlled experiments across multiple models, capturing metrics like success rate and average response time for reproducible results.
Offer LLM Tester as a standalone tool for developers or QA teams, with monthly subscriptions for access to advanced features like custom report templates and CI/CD integration.
Provide a consulting service where customers send their test samples and prompts, and the company runs the benchmarks and delivers analysis reports.
Embed LLM Tester into a larger AI development platform (e.g., model marketplace or MLOps tool) as a premium feature that helps users select the best model for their needs.
💬 Integration Tip
Integrate LLM Tester into your CI/CD pipeline by running it as a step in your build process; the JSON report can be parsed by automated quality gates to flag regressions.
Scored May 23, 2026
Real-time search engine supporting web search, vertical domain search, parallel batch search, and URL content extraction.
Manage Feishu (Lark) calendars by listing, searching, checking schedules, syncing events, and marking tasks with automated date extraction.
Process multiple items with progress tracking, checkpointing, and failure recovery.
cad reference tool
Search, install, and create OpenClaw skills using intelligent matching across built-in, local, and GitHub skill repositories.
Use when building CLI tools, implementing argument parsing, or adding interactive prompts. Invoke for CLI design, argument parsing, interactive prompts, progress indicators, shell completions.